Skip to content
Skip to content

Independent e-magazine

the OUTSPOKEN digest

Google's New Gemini Avatar Talks Back With a Face, Lip-Syncing Live Across 97 Languages

Gemini 3.8 Live with Live Avatar pairs real-time dialogue with near-instant video generation, letting an animated persona hold a conversation, switch languages mid-sentence, and keep working while it talks.

Outspoken Digest Technology Desk

Saturday, September 26, 2026/2 min read

Google's Gradient Canopy building on its Mountain View, California campus, photographed in April 2025
Photo: GualdimG via Wikimedia Commons (CC BY-SA 4.0)

Google made its newest Gemini avatar generally available inside Gemini Enterprise this week, pairing the company's live-dialogue models with near-real-time video generation so that the assistant does not just answer in text or speech but appears to answer, with a face, expressions and lip movements that track what it is saying as it says it. Google's own announcement describes Gemini 3.8 Live with Live Avatar as combining its live conversational model with a video layer that renders convincingly enough to sustain eye contact and natural pauses through an entire exchange.

What makes it different from a video call with a bot

The headline feature is language switching. The Tech Portal's report says the avatar can move between any of 97 supported languages mid-conversation, with lip-sync and facial expression adjusting on the fly, without the visual stutter or drift that has made earlier AI-generated video feel obviously synthetic. Google's demo shows the avatar checking a hotel guest in while a booking system query runs in the background, the face continuing to talk and react rather than freezing while the tool call completes, a detail meant to show the system handling backend work without breaking the illusion of a continuous conversation.

Businesses can build their own face

Companies using Gemini Enterprise can select from a set of ready-made avatars or build a custom one from a single high-quality reference photograph, a feature Google is restricting for now to enterprise customers on an allowlist, a deliberate throttle on a capability that could otherwise be used to generate a convincing likeness of a real person from one image. Every piece of audio and video the system produces carries Google's SynthID watermark, invisible to viewers but detectable by software, intended to let the output be identified as AI-generated after the fact even when it is not obvious on first viewing.

Part of a broader shift toward agents with faces

The launch lands in a week when AI companies have been racing to make their assistants feel less like chat windows and more like presences, whether that is Meta's Muse shopping agent running its own browser or OpenAI's research agents operating with enough autonomy to overstep their intended scope. A talking, lip-syncing avatar that keeps working in the background while it holds a conversation is a step toward assistants that behave like colleagues rather than tools, and Google's decision to gate custom avatars behind an allowlist suggests the company is aware of how quickly that same capability could be misused outside a controlled enterprise setting.

Published in The Outspoken Digest

Editorial desk

Outspoken Digest Technology Desk

Software, hardware, artificial intelligence and what they change for everyone else.

Newsletter

The Digest, in your inbox

One edition, sent when it is ready. No noise, and your address is never passed on.

We send a confirmation first. One click to leave, always.

Share this story

the OUTSPOKEN digest

Beyond boundaries. Independent stories on technology, culture, and the trends shaping how we live.