study
Reviews call Kling 3.0 the best for lip-sync. It generated zero audio for us.
Four dialogue prompts, four silent clips. The mouths move on cue; the sound never shows up.
Search for a review of Kling 3.0 and you will find a version of this sentence on more than one site: Kling is best-in-class for realistic human faces and movement, with impressive lip-syncing that makes it a strong choice for dialogue-driven video. We ran it. It does not generate audio at all.
What we tested
Our battery includes four lip-sync prompts specifically designed to test spoken dialogue: a direct-address line, a phrase spoken through a beard, a two-person exchange with turn-taking, and a sung line with visible breath. Every prompt asks the model to produce a person saying specific words, on camera.
We ran all four through Kling 3.0 via the platform we use for generation (Magnific), then had a vision model transcribe whatever audio was present.
The result: silence, four times
| Prompt | What happened |
|---|---|
| Direct address | Lips move correctly through the full line. No audio track at all. |
| Bearded speaker | Same — mouth shapes are there, sound is not. |
| Two-person exchange | Correct turn-taking, one person's mouth moves while the other stays still. Still silent. |
| Sung line | Breath vapour and mouth shapes for the sung phrase. No vocal track. |
This is not a sync problem — the mouth movements are genuinely well-timed to where speech would land. There is simply no sound. Watch any of the four clips linked above with the volume up; that is the whole finding, no interpretation needed.
Why this matters more than a bad score would
A model that generates audio badly is a quality problem you can work around — dub over it, regenerate, adjust the prompt. A model that does not generate audio at all, while being recommended across the web specifically for its dialogue capability, is a capability gap that will not show up until you are already relying on it in production.
We initially scored these clips as failures and started writing a lower audio number into the record. Then we caught our own mistake: scoring a missing capability as a "0" implies the model tried and failed, which is different from a model that was never asked to try. We corrected our own policy mid-benchmark — a feature a model does not offer gets left unscored, not zeroed. Kling's model page now marks native audio as unavailable, corrected from an earlier listing that assumed the market description was accurate.
A likely explanation
Kling's family includes a separate "Omni" variant that is genuinely built for multi-language dialogue on a shared audio timeline. Our reading is that market coverage describing Kling 3.0 as strong for dialogue is describing that variant, or an access path we have not tested, not the base model available through the platform we used. Either way: the base Kling 3.0, as we could access it, is silent.
The takeaway
If dialogue or generated sound is part of your workflow, do not take a review's word for it — including ours, for anything we have not personally tested. Watch a real clip with the volume on before you commit to a model. We do this so you do not have to, and this is exactly the kind of thing that only shows up when someone actually runs the test.
Related: Kling 3.0
Jose Florido — Runs the SlopTV benchmark. Every score on this site comes from a generation he paid for and a clip you can watch.