Google Gemini Omni
What it is actually for
Gemini Omni landed on 19 May 2026 as a multimodal model rather than a video model, with video generation as one capability among several. Reporting at the time put it ahead of Sora 2 on the capabilities Sora users cared about most — which, given Sora was already on its way out, made it a natural landing spot.
Why it is listed separately from Veo
Because they are different products with different shapes. Veo 3.1 is a dedicated video model you call to make a clip. Omni is a general multimodal system where video is one output among several, which matters if your pipeline is doing more than one kind of generation and you would rather not stitch together three vendors.
What we do not know yet
Almost everything, honestly. It is the newest entry we track and we have not put it through the battery. We have listed it because leaving a relevant model off a comparison is worse than listing it with an explicit gap where the numbers should be.
Capabilities
| Native audio | Yes |
| Multi-shot sequences | No |
| Image to video | Yes |
| Public API | Yes |
Benchmark scores
This model has not been through a scored round yet.