Google’s Gemini Omni, accessed through the new Google Vids interface, was put through a hands-on lab test to see if the AI could keep a single subject-courier stable after three successive video edits.

Why the test matters

Generative video tools promise anyone can swap backgrounds, add lighting or sprinkle effects with a single text command. For a creator the promise is simple: start with a raw clip, type “make it rain at night,” and walk away with a polished short. Most users don’t see the hidden cost of “drift”—the gradual loss of the original subject’s appearance, pose or motion as the AI interprets each new instruction. If a courier’s face subtly changes after the second edit, the final piece looks unprofessional and may need costly manual clean-up.

The experiment

Baseline clip – A static-camera studio shot of one person wearing a navy windbreaker, holding a small cardboard box in the left hand, and taking two steps toward the camera. The lighting was even, the background plain, and the motion simple.

Edit sequence – Three edits were applied one after another:

  1. Background swap – Replace the studio with a rainy city street at night.
  2. Lighting tweak – Add a blue rim light around the courier and a warm glow from a shop window.
  3. Atmosphere layer – Sprinkle light raindrops and a thin veil of steam.

A “direct control” run paired all three instructions in a single prompt, letting the AI generate the final clip in one go.

How consistency was scored

Twelve criteria formed the scorecard, grouped into four themes:

  • Character stability – Face, hair and clothing remain identical.
  • Scene stability – Camera angle and subject position stay fixed.
  • Edit precision – The AI follows the instruction without altering anything else.
  • Temporal quality – No flicker, warping or jitter appears during movement.

Each clip was evaluated at the first frame, the midpoint of the two-step walk, and the final frame. The scoring revealed a clear pattern.

What the numbers say

When edits were stacked, the AI was evaluated using the scorecard. The single-prompt control received the same evaluation.

Who wins and who loses

Studios that can afford a single, comprehensive prompt gain a smoother workflow. Developers of generative video systems face a clear engineering hurdle—the need to keep a stable latent representation of the subject across sequential transformations.

Counter-argument

Some users argue incremental editing offers more creative control. By adjusting one element at a time, they can fine-tune each layer before moving on. The test shows that this flexibility costs visual fidelity. If a creator accepts minor subject changes for granular control, the current Gemini Omni workflow may still be viable.

What to watch next

  • Model refinements that lock the subject’s latent code across prompts.
  • User-feedback loops allowing creators to flag drift and request a “re-anchor” of the subject.

Takeaway

Gemini Omni can produce a polished, multi-element video when fed a complete set of instructions at once, but its subject fidelity erodes when edits are stacked.