SlopTV

study

We tested AI video hands. One model grew a third arm.

Four hand-specific prompts, three models, and the exact moment each one breaks.

By Jose Florido ยท

We ran the same 4 hand-specific prompts through Sora, Kling and Grok. Sora went 4-for-4 clean. Kling and Grok both produced genuine polydactyly on an open-palm test, and Grok hallucinated a third hand entirely on a knife-and-pepper prompt. Both failures clustered on unconstrained motion, not anchored contact.

Everyone in this field knows AI video struggles with hands. Nobody quantifies it with the same prompts across models, and almost nobody shows you the actual clip. We ran the same 4 hand-specific prompts through Sora 2, Kling 3.0 and Grok, and the failures ranged from a barely-visible flaw to a clip that hallucinated a third arm.

Why hands specifically break

Hands pack more structural detail into a small area of frame than almost anything else models generate: five digits, each with multiple joints bending on specific axes, frequently overlapping and occluding each other. Training footage also tends to show hands smaller and less clearly than faces. The result is a part of the human body that has to be gotten exactly right โ€” digit count, joint angles, overlap order, all simultaneously โ€” with less clean training signal than everything else in frame.

That is the theory. Here is what it actually looks like across three current models.

The four tests

TestSora 2Kling 3.0Grok
Tying shoelaces5.004.174.17
Fist to open palm4.833.172.50
Fretting a guitar chord5.005.004.50
Chopping a pepper5.002.832.67

Sora went 4-for-4 clean. The other two only failed on the two hardest prompts — the ones asking for a specific finger count and for grip-plus-repeated-motion-plus-an-object-changing-state.

Failure 1: counting fingers, both models add one

The prompt asks for a hand opening from a closed fist to a fully spread palm — a countable, verifiable action. Kling and Grok both produced genuine polydactyly: a sixth digit appears partway through the opening motion, with visible morphing where the extra finger emerges. Sora kept five fingers throughout, losing only a small amount of fidelity to a hard cut mid-motion.

Failure 2: chopping vegetables, and Grok's third hand

This is the worst single failure in our entire hands category. Asked for hands slicing a pepper with a knife, Grok generated three separate hands manipulating the knife and pepper at once — not warped fingers, an extra limb. Kling's version was less extreme but still broke: finger warping on the knife grip, and pepper slices that duplicate and slide across the board without real cutting resistance. Watch the clip yourself; it is easier to see than to describe.

What held up: contact-anchored motion

The two prompts every model handled well — tying a boot lace, fretting a guitar — both anchor the hand against a physical object for the whole shot. The two that broke — an open palm in empty space, a knife-and-pepper interaction with changing state — both involve unconstrained motion or an object that has to change shape correctly. That lines up with the standard advice for working around this: never let hands move freely with nothing to anchor against.

The takeaway

"AI video struggles with hands" is true but useless for deciding anything. Which model, on which kind of shot, failing in which specific way, is the version you can actually plan around. Right now: if your shot needs a hand doing something with nothing to hold, or manipulating an object that changes shape, budget for retries with Kling or Grok, or reach for Sora while it is still available. We will keep re-running this exact battery as new models enter the leaderboard.


Jose Florido โ€” Runs the SlopTV benchmark. Every score on this site comes from a generation he paid for and a clip you can watch.