review
Gemini Omni Review: What Our Own 40-Prompt Benchmark Found
Gemini Omni scored 4.91/5 in our testing, the highest of any model we have run, and it is the only one that produced real synced audio on every lip-sync prompt.
Gemini Omni is now the top model in our benchmark, and it got there on the back of a category every other model we have tested fails outright: lip-sync. It is the only model in our round that produces real spoken audio matching the requested line, synced to the mouth on screen, instead of a silent clip or garbled ambient noise.
What we actually measured
We ran Gemini Omni through our fixed battery of 40 prompts across 10 categories, scored blind against the same rubric we use for every model. It came out at 4.91/5 overall, the highest score we have recorded, ahead of Kling 3.0 Turbo at 4.89.
| Dimension | Score |
|---|---|
| Prompt fidelity | 4.92/5 |
| Visual quality | 4.94/5 |
| Physics and motion | 4.92/5 |
| Subject and scene consistency | 4.95/5 |
| Audio and sync | 4.82/5 |
Every dimension sits at 4.8 or above. Nothing in our data is a weak point in the way hands or lip-sync are weak points for most models we have tested; the gap between Gemini Omni's best and worst category is under a quarter of a point.
By category
| Category | Score |
|---|---|
| Lip-sync and dialogue | 4.99/5 |
| Human faces | 4.96/5 |
| Multi-shot continuity | 4.96/5 |
| Camera movement | 4.96/5 |
| Style consistency | 4.96/5 |
| Image to video | 4.96/5 |
| Hands and fingers | 4.91/5 |
| On-screen text | 4.91/5 |
| Liquid physics | 4.79/5 |
| Animal anatomy | 4.73/5 |
Liquids and animals are the lowest-scoring categories, and even those are middling rather than broken. The condensation on a melting-ice-cube prompt formed faster than physically plausible, a galloping horse showed some foot-sliding and a stiffer torso than a real gait would have, and a cat jumping onto a counter had a slightly floaty apex mid-leap. None of these were the kind of outright failure we have flagged in other models' runs, like an animal never leaving the ground or a hand losing a finger.
The category every other model fails
Lip-sync is the highest-scoring category in our data for Gemini Omni, and it is worth explaining why that is unusual. Every other model in our round either generates a silent clip for a spoken-line prompt, or animates a mouth moving with no audio behind it, or in Kling 3.0's case, produces zero audio at all despite third-party marketing describing native dialogue.
Gemini Omni generated real spoken audio for all four of our lip-sync prompts, and the words matched what we asked for. A woman delivered our exact test line to camera with her mouth movements synced to the audio. A two-speaker exchange had both people say their assigned lines with natural back-and-forth timing. A singing prompt produced sustained, on-pitch vowel sounds synced to visible mouth shapes, not just a mouth flapping to a beat. We confirmed the audio track itself was real, not silent, on all 40 clips by checking the actual waveform, not just trusting a metadata flag.
This is the first time in our benchmark that a model has cleared this bar. It is also, on its own, the reason Gemini Omni took the top spot: strip the audio dimension out and its overall score is close enough to Kling 3.0 Turbo's that the two would be within each other's margin of error.
What else stood out
- The 360-degree camera orbit completed for the first time in our testing. Every other model we have run either stops partway around the subject or breaks continuity mid-turn. Gemini Omni's orbit prompt passed behind the subject and returned to the front.
- The product-rotation prompt also completed a full turn, with the mug's handle disappearing and reappearing on the correct side, something none of the other models in our round managed cleanly.
- The multi-shot day-to-dusk prompt actually changed the lighting. Every other model we have tested on this exact prompt renders one continuous shot or a static frame instead of the requested lighting change between cuts.
Who should use it, and what to check first
- Anything that needs a character speaking on camera: this is the only model in our testing that reliably produces matching audio and lip movement. If dialogue is the point of the clip, this is currently the model to test first.
- Camera moves that need to complete, like a full orbit or a full product turn: strongest result we have measured for either.
- Close-up liquid physics or an animal as the hero of the shot: still the two weakest categories in our data for this model, though not by a wide margin. Worth a test clip before committing if the shot depends on either.
Full clips, per-prompt scores and every category breakdown are on the Gemini Omni model page. The lip-sync leaderboard shows how every model we have tested compares on audio specifically. Methodology, including how we score and verify findings before publishing them, is on the methodology page.
Related: Google Gemini Omni
Mara Voss: Runs the prompt battery and scores every clip before it is published. Came to AI video testing from QA automation, where she spent years building test suites designed to break things on purpose.