Lead Benchmark Engineer
Mara Voss
Runs the prompt battery and scores every clip before it is published. Came to AI video testing from QA automation, where she spent years building test suites designed to break things on purpose.
Published on SlopTV
Kling 3.0 vs Grok: same 40 prompts, category by category
Kling 3.0 and Grok score within 0.22 of each other overall. The category breakdown tells a very different story: Kling wins on style and faces, Grok sweeps camera moves, animals and lip-sync outright.
comparison 12 Sep 2026
We tested AI video hands. One model grew a third arm.
Sora, Kling and Grok on the same 4 hand prompts. Two models gave a test subject six fingers. One generated three separate hands on a single arm. Here is exactly where and why.
study 12 Sep 2026
Reviews call Kling 3.0 the best for lip-sync. It generated zero audio for us.
Kling 3.0 is widely described as best-in-class for AI dialogue and lip-sync. We ran four spoken-line prompts through it and got perfectly-timed silent mouth movement, zero audio, every time.
study 12 Sep 2026