Google’s Gemini Omni, accessed through the new Google Vids interface, was put through a hands-on lab test to see if the AI could keep a single subject-courier stable after three successive video edits.
Why the test matters
Generative video tools promise anyone can swap backgrounds, add lighting or sprinkle effects with a single text command. For a creator the promise is simple: start with a raw clip, type “make it rain at night,” and walk away with a polished short. Most users don’t see the hidden cost of “drift”—the gradual loss of the original subject’s appearance, pose or motion as the AI interprets each new instruction. If a courier’s face subtly changes after the second edit, the final piece looks unprofessional and may need costly manual clean-up.
The experiment
Baseline clip – A static-camera studio shot of one person wearing a navy windbreaker, holding a small cardboard box in the left hand, and taking two steps toward the camera. The lighting was even, the background plain, and the motion simple.
Edit sequence – Three edits were applied one after another:
- Background swap – Replace the studio with a rainy city street at night.
- Lighting tweak – Add a blue rim light around the courier and a warm glow from a shop window.
- Atmosphere layer – Sprinkle light raindrops and a thin veil of steam.
A “direct control” run paired all three instructions in a single prompt, letting the AI generate the final clip in one go.
How consistency was scored
Twelve criteria formed the scorecard, grouped into four themes:
- Character stability – Face, hair and clothing remain identical.
- Scene stability – Camera angle and subject position stay fixed.
- Edit precision – The AI follows the instruction without altering anything else.
- Temporal quality – No flicker, warping or jitter appears during movement.
Each clip was evaluated at the first frame, the midpoint of the two-step walk, and the final frame. The scoring revealed a clear pattern.
What the numbers say
When edits were stacked, the AI was evaluated using the scorecard. The single-prompt control received the same evaluation.
Who wins and who loses
Studios that can afford a single, comprehensive prompt gain a smoother workflow. Developers of generative video systems face a clear engineering hurdle—the need to keep a stable latent representation of the subject across sequential transformations.
Counter-argument
Some users argue incremental editing offers more creative control. By adjusting one element at a time, they can fine-tune each layer before moving on. The test shows that this flexibility costs visual fidelity. If a creator accepts minor subject changes for granular control, the current Gemini Omni workflow may still be viable.
What to watch next
- Model refinements that lock the subject’s latent code across prompts.
- User-feedback loops allowing creators to flag drift and request a “re-anchor” of the subject.
Takeaway
Gemini Omni can produce a polished, multi-element video when fed a complete set of instructions at once, but its subject fidelity erodes when edits are stacked.
