news
Google ships Live Avatar for Gemini 3.8 Live, with 97 languages and no comparative quality numbers
The Gemini Enterprise launch pairs near real-time video generation with speech, watermarks every stream with SynthID, and gates custom avatars behind an allowlist, without publishing a single benchmark.
Google's newest video generation product shipped this week with a language count, a watermarking policy and an allowlist, and not one number describing output quality. Gemini 3.8 Live with Live Avatar became generally available in Gemini Enterprise, pairing near real-time video generation with the company's native speech-to-speech model so an agent can hold a conversation with a visible face attached.
The product is a talking face that runs alongside tool calls
Google describes the feature as creating "an experience that listens, sees, and speaks with a dynamic visual persona," with lip-syncing, facial expressions and turn-taking handled as part of the same stream rather than bolted on after the fact. The avatar processes visual and audio input at the same time, so camera feeds and screen sharing come in while the video face is being generated. Asynchronous tool calling lets the agent fetch data in the background and keep talking while it waits, which is the part that separates this from a conventional avatar video renderer.
97 languages is the only hard figure in the launch
The one measurable claim Google puts forward is multilingual coverage: native speech-to-speech synchronization across 97 languages, with lip-sync and expressions adapting when a conversation switches language mid-stream, and no visual drift during the transition. That last part is a quality claim stated as an assertion, not a measurement. There is no side-by-side against other avatar or video systems, no human preference test, no error rate on lip-sync accuracy, and no published figure for how far "near real-time" actually is in milliseconds. For a launch whose entire value proposition is perceived realism in a live conversation, that is a conspicuous gap, and it is the same gap that shows up across the video model releases we track, from Veo 3.1 to Gemini Omni.
Identity controls are the part Google did specify
The deployment rules are unusually concrete by comparison. Customers can pull from a library of curated, pre-built avatars, while creating a custom avatar from a reference photo and audio sample sits behind an enterprise allowlisting and verification process. Every generated audio and video stream carries an imperceptible SynthID watermark, which Google frames as making the output verifiable after the fact. Availability is limited to US and EU endpoints with provisioned throughput and enterprise data governance, and the related Gemini 3.8 Live Extended Thinking variant remains in private preview.
Why this one is genuinely hard to score
Standard video benchmarks assume a prompt goes in and a finished clip comes out, which is not how a live avatar works. The output depends on what the user says, how fast they interrupt, which language they switch into and how long the session runs, so a fixed prompt set cannot reproduce it. Evaluating it properly means scoring a conversation, not a file, and nobody in this space has published a shared method for doing that. Until someone does, claims about lip-sync fidelity and expression naturalness in live avatars will stay exactly where Google left them, as adjectives.
Sources: Google, Google Cloud, Engadget.
Priya Shenoy: Tracks what AI video actually costs across the platforms that resell access to the same handful of models. Treats a pricing page as a claim, not a fact, until someone checks it.