Gemini 3.8 Live with Live Avatar can generate synchronized video of a speaking AI character while the underlying system listens, takes visual input, talks to the user and calls external tools in the background.
The obvious attraction is the avatar. The more useful question is what sits behind it.
Google is combining live speech, visual understanding, background tool execution and synchronized video inside the same interactive session. That could make customer-service agents, digital concierges, onboarding assistants and guided support systems feel considerably more natural.
The limits are just as important.
Live Avatar is aimed at businesses building on Google’s enterprise AI stack. Custom avatars are gated. Google’s own model card says continuous avatar interactions currently last a few minutes rather than extended hours.
So the launch is less "digital human joins every video call" and more "Google has given short, task-focused AI agents a much more human interface."
Choosely has not tested Gemini 3.8 Live Avatar hands-on. This early assessment is based on Google’s launch material, technical documentation, model card and published pricing.
Google has given its live agents a face
Google announced Gemini 3.8 Live with Live Avatar on September 24, nine days after introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking.
The September 24 release added Live Avatar video output and moved Gemini 3.8 Live into general availability in Gemini Enterprise, with US and EU endpoints, provisioned throughput and enterprise deployment controls.
Google had first previewed Live Avatar at Google Cloud Next 2026.
Several of the capabilities underneath the avatar arrived with Gemini 3.8 Live itself, including background tool calling, live visual understanding and multilingual conversation.
Live Avatar adds synchronized video to that foundation.
Google’s developer documentation specifies avatar output at 24 frames per second. The generated facial movement and lip synchronization follow the model’s synthesized speech.
Developers can use Google’s prebuilt avatars or, with additional approval, create a custom avatar from a reference image.
That changes the interface more than the intelligence underneath it.
The agent was already able to listen, see and act. Now the user can look back at something while it does those things.
The first limitation is surprisingly basic: time
The polished demos can make Live Avatar look like an always-on digital person.
Google’s own model card puts a much clearer boundary around that idea.
It says Gemini 3.8 Live with Live Avatar currently supports a few minutes of continuous interaction rather than extended hours.
That immediately narrows the most realistic use cases.
A hotel check-in, guided product setup, claims intake, short support interaction or concierge request fits comfortably inside that window.
A persistent digital employee expected to sit in meetings, teach a long class or remain continuously available for hours does not.
That distinction matters because "AI avatar" can describe two very different products.
One is a visible interface for completing a bounded task.
The other is an artificial person expected to maintain a relationship or role over time.
Google has shipped something much closer to the first category.
Teams evaluating Live Avatar should treat session length as a product constraint, not a detail buried in a model card.
Tool use is what makes this more than an avatar demo
Gemini 3.8 Live supports asynchronous function calling.
An agent can trigger an API or backend action while the conversation continues instead of forcing every external task to create an awkward pause.
Google demonstrates the capability with a hotel check-in flow where the agent keeps talking while a tool call runs in the background.
That is a more consequential capability than better lip synchronization.
A customer-service agent becomes genuinely useful when it can keep the interaction moving while checking an account, retrieving information, updating a booking or triggering another business system.
The avatar improves the interface.
The external actions determine whether anything useful actually gets done.
This also raises the standard for reliability.
If an agent looks like an attentive human representative but sends the wrong request to a CRM or confidently gives the wrong account information, the visual polish quickly becomes irrelevant.
The better the interface becomes, the less forgiving the underlying agent can afford to be.
Vision, voice and the face now sit in one session
Gemini 3.8 Live accepts live visual input alongside audio.
Google says developers can send camera feeds and screen shares into the same session.
That creates interactions where voice alone would be awkward.
A support agent could react to something a customer is showing through their camera. A software assistant could discuss what is visible on a shared screen. An onboarding agent could respond to the user’s current step rather than relying entirely on a verbal description.
Google has also demonstrated a claims-intake scenario where a user shows physical damage through a camera while the agent processes information in the background.
The model can automatically detect and transition across 97 supported languages during a conversation, according to Google.
For Live Avatar, Google says lip synchronization and facial expressions adapt as the language changes without degrading video fidelity or creating visible drift.
Choosely has not independently tested that claim.
If it survives real conditions such as mixed-language speech, accents, interruptions and noisy environments, the combination could be particularly useful for international support and hospitality deployments.
The important shift is the consolidation.
Speech, visual context, tool execution and the visible representative no longer need to feel like separate layers stitched together around the user.
Custom avatars come with much tighter controls
Google provides a library of prebuilt avatars that can be paired with supported voices.
Custom avatars are possible too.
Developers can provide a reference image and have Gemini generate a responsive avatar from that likeness.
But this is not an unrestricted self-service cloning tool.
Custom avatar creation is currently available only to selected customers through an enterprise allowlisting and verification process.
Google also puts responsibility for likeness rights on the customer. Businesses must secure the necessary consent and rights for any face or voice samples they provide.
Its reference-image guidance explicitly says not to use images of minors, celebrities or offensive content.
Google says Live Avatar audio and video output also carries SynthID watermarking.
SynthID embeds an imperceptible signal into generated media that can help supported detection systems identify Google AI output.
That is useful, but it should not be confused with disclosure.
A watermark does not automatically tell the person on the other side of the interaction that they are talking to AI.
Visible or spoken disclosure remains a deployment decision for the business using the avatar.
As generated representatives become more convincing, that distinction becomes harder to treat as a secondary product detail.
The $1-per-million-token headline hides the real video cost
There is no simple consumer subscription price for Live Avatar.
Google Cloud charges Gemini 3.8 Live API usage by modality on its non-global endpoints.
The current standard rates are:
- text input: US$0.75 per 1 million tokens
- audio input: US$3 per 1 million tokens
- image or video input: US$1 per 1 million tokens
- text output: US$4.50 per 1 million tokens
- audio output: US$12 per 1 million tokens
- avatar video output: US$1 per 1 million tokens
At first glance, avatar video looks like one of the cheapest lines on the rate card.
Token conversion changes that picture.
Google converts avatar video output at 6,192 tokens per second and only charges that video output while the avatar is speaking.
At US$1 per million tokens, the avatar video layer alone works out to roughly US$0.37 per minute of speaking time.
Audio output is converted at 25 tokens per second. At US$12 per million tokens, that adds roughly US$0.018 per minute.
Input audio, camera or image input, text, tool use and other infrastructure costs sit on top.
Live API sessions have another cost characteristic worth understanding.
Google bills each turn using the tokens currently held in the session context. Previous conversation tokens are therefore reprocessed and charged again as later turns occur.
Even inside the current short-session limit, cost can rise as the conversation accumulates context.
For an enterprise customer, that does not automatically make Live Avatar expensive.
It does mean the apparently simple "$1 per million avatar tokens" number is a poor way to understand what the visible agent actually costs to run.
This is not a new Gemini personality on your phone
This is the consumer confusion most likely to follow the launch.
Gemini 3.8 Live itself already reaches consumer products.
Google has rolled it into Search Live for general users, while Gemini 3.8 Live Extended Thinking is available through Gemini Live with additional integrations across supported Google Workspace products and plans.
That does not mean Live Avatar has suddenly appeared as a visual Gemini personality inside the normal Gemini app.
The synchronized avatar capability announced on September 24 belongs to Google’s enterprise and developer agent stack.
For an ordinary Gemini user, there is very little to change today.
For a company already building voice or conversational agents, the decision is much more immediate.
That distinction also explains why the launch can look bigger on social media than it feels to most Gemini users.
The technology is visually striking.
The addressable user today is considerably narrower.
Google just moved deeper into the real-time avatar market
Gemini Live Avatar overlaps with the broader AI-avatar market, but the competitive pressure is not evenly distributed.
Many avatar products focus on completed media: presenter videos, translated training content, marketing clips or social assets.
That remains a different job.
Gemini Live Avatar is built around interaction.
The user can talk to it. The agent can see live context. It can trigger external systems. The conversation can continue while those actions run.
That puts more direct pressure on the emerging category of real-time conversational avatars and digital agents.
Google also gains an obvious advantage from owning more of the stack.
The model, voice, visual understanding, agent tooling and avatar can now come from the same provider.
That can simplify deployment for businesses already committed to Google Cloud.
It does not automatically make Google the best avatar platform.
Avatar quality, character control, latency, session duration, workflow tooling and customization all still matter. Google’s current few-minute interaction ceiling alone leaves obvious territory for specialists designed around longer conversations.
A company primarily producing prerecorded avatar videos should not assume this launch changes its tool choice.
A company building a live support, concierge or onboarding agent should pay much closer attention.
The practical buyer question is whether the face improves the task
Teams can test Google’s prebuilt avatars through Agent Platform’s Stream realtime interface without entering the custom-avatar process.
That is probably where the evaluation should start.
Pick a short interaction with a clear beginning and end.
A check-in. Claims intake. Guided setup. Product walkthrough. Support task.
Then measure whether the visible avatar actually improves comprehension, trust, task completion or customer experience.
Do not start by trying to design an all-day virtual employee around a feature Google currently describes in minutes.
The face is only valuable when it improves the job the agent was already supposed to do.
Choosely Verdict
Gemini 3.8 Live Avatar matters because Google has brought live speech, visual understanding, tool execution and synchronized avatar video into the same agent platform.
The launch makes AI agents easier to present as visible participants rather than disembodied voices.
That could be valuable in short, customer-facing interactions where eye contact, expression and visual context genuinely improve the experience.
The limits keep the announcement grounded.
Continuous avatar sessions currently run for minutes rather than hours. Custom avatars are gated. The video layer costs materially more per minute than its headline token price suggests. Choosely has not independently tested Google’s avatar quality, latency or multilingual claims.
There is not enough evidence to call Gemini Live Avatar a new default for AI avatars.
There is enough to say Google has made short-form visible conversational agents significantly easier to build inside its own AI stack.
For businesses already deploying voice agents, that is worth evaluating.
For everyone else, the demo is currently more advanced than the need.
The Change Brief
Get the week’s AI changes in one clear read
Pricing moves, tool launches, free-tier changes and practical stack updates, filtered for people who actually use these tools.
Stay ahead of AI without following it all day. We’ll send you what matters each week.
Continue reading
Related reads
AI Tool Recommendations
ChatGPT Voice vs Claude Voice: Which One Should You Actually Talk To?
ChatGPT Voice is the stronger documented fit for fluid everyday conversation, while Claude Voice stands out for model choice and a more deliberate turn-taking style.
AI Strategy
Gemini Hacked Three Companies. Did Google's AI Actually Go Rogue?
Gemini reached three real companies during a cyber test, but the evidence points to failed containment, agent overreach and uneven model behavior rather than a simple rogue-AI story.
AI Strategy
MCP Explained: How AI Agents Connect to Your Tools and What They Can Access
MCP is the protocol many AI agents use to connect to tools, files, apps, and workflows. Here is what it is, how it works, and what permissions to review before you enable it.
Product Update
Claude's Invisible Watermark Is Real. The Panic Is Ahead of the Facts
Anthropic's invisible Claude watermarks are real, but claims that every current response and document is already secretly tracked go beyond the evidence.