SlopTV

review

Gemini Omni Review: What Our Own 40-Prompt Benchmark Found

Gemini Omni scored 4.91/5 in our testing, the highest of any model we have run, and it is the only one that produced real synced audio on every lip-sync prompt.

By Mara Voss ยท

Gemini Omni scored 4.91/5 in our own 40-prompt benchmark, the highest of any model tested, ahead of Kling 3.0 Turbo at 4.89. It is the only model in our round that generates real spoken audio synced to on-screen lip movement, matching the requested dialogue rather than staying silent or producing mismatched mouth motion. Its weakest categories, liquid physics and animal anatomy, still scored 4.7 or higher.

Gemini Omni is now the top model in our benchmark, and it got there on the back of a category every other model we have tested fails outright: lip-sync. It is the only model in our round that produces real spoken audio matching the requested line, synced to the mouth on screen, instead of a silent clip or garbled ambient noise.

What we actually measured

We ran Gemini Omni through our fixed battery of 40 prompts across 10 categories, scored blind against the same rubric we use for every model. It came out at 4.91/5 overall, the highest score we have recorded, ahead of Kling 3.0 Turbo at 4.89.

DimensionScore
Prompt fidelity4.92/5
Visual quality4.94/5
Physics and motion4.92/5
Subject and scene consistency4.95/5
Audio and sync4.82/5

Every dimension sits at 4.8 or above. Nothing in our data is a weak point in the way hands or lip-sync are weak points for most models we have tested; the gap between Gemini Omni's best and worst category is under a quarter of a point.

By category

CategoryScore
Lip-sync and dialogue4.99/5
Human faces4.96/5
Multi-shot continuity4.96/5
Camera movement4.96/5
Style consistency4.96/5
Image to video4.96/5
Hands and fingers4.91/5
On-screen text4.91/5
Liquid physics4.79/5
Animal anatomy4.73/5

Liquids and animals are the lowest-scoring categories, and even those are middling rather than broken. The condensation on a melting-ice-cube prompt formed faster than physically plausible, a galloping horse showed some foot-sliding and a stiffer torso than a real gait would have, and a cat jumping onto a counter had a slightly floaty apex mid-leap. None of these were the kind of outright failure we have flagged in other models' runs, like an animal never leaving the ground or a hand losing a finger.

The category every other model fails

Lip-sync is the highest-scoring category in our data for Gemini Omni, and it is worth explaining why that is unusual. Every other model in our round either generates a silent clip for a spoken-line prompt, or animates a mouth moving with no audio behind it, or in Kling 3.0's case, produces zero audio at all despite third-party marketing describing native dialogue.

Gemini Omni generated real spoken audio for all four of our lip-sync prompts, and the words matched what we asked for. A woman delivered our exact test line to camera with her mouth movements synced to the audio. A two-speaker exchange had both people say their assigned lines with natural back-and-forth timing. A singing prompt produced sustained, on-pitch vowel sounds synced to visible mouth shapes, not just a mouth flapping to a beat. We confirmed the audio track itself was real, not silent, on all 40 clips by checking the actual waveform, not just trusting a metadata flag.

This is the first time in our benchmark that a model has cleared this bar. It is also, on its own, the reason Gemini Omni took the top spot: strip the audio dimension out and its overall score is close enough to Kling 3.0 Turbo's that the two would be within each other's margin of error.

What else stood out

  • The 360-degree camera orbit completed for the first time in our testing. Every other model we have run either stops partway around the subject or breaks continuity mid-turn. Gemini Omni's orbit prompt passed behind the subject and returned to the front.
  • The product-rotation prompt also completed a full turn, with the mug's handle disappearing and reappearing on the correct side, something none of the other models in our round managed cleanly.
  • The multi-shot day-to-dusk prompt actually changed the lighting. Every other model we have tested on this exact prompt renders one continuous shot or a static frame instead of the requested lighting change between cuts.

Who should use it, and what to check first

  • Anything that needs a character speaking on camera: this is the only model in our testing that reliably produces matching audio and lip movement. If dialogue is the point of the clip, this is currently the model to test first.
  • Camera moves that need to complete, like a full orbit or a full product turn: strongest result we have measured for either.
  • Close-up liquid physics or an animal as the hero of the shot: still the two weakest categories in our data for this model, though not by a wide margin. Worth a test clip before committing if the shot depends on either.

Full clips, per-prompt scores and every category breakdown are on the Gemini Omni model page. The lip-sync leaderboard shows how every model we have tested compares on audio specifically. Methodology, including how we score and verify findings before publishing them, is on the methodology page.

Some links on this page are affiliate links. See affiliate disclosure for exactly what that does and does not change about what we publish.

Related: Google Gemini Omni


MV

Mara Voss: Runs the prompt battery and scores every clip before it is published. Came to AI video testing from QA automation, where she spent years building test suites designed to break things on purpose.