A face for a live conversation
Google has introduced Live Avatar for Gemini 3.8 Live, adding generated visual presence to its conversational model. The Google announcement, published on 24 September, says the feature pairs live dialogue with low-latency streaming video. It is available in Gemini Enterprise. This is a new visual capability built on last week’s Live model release, not another announcement of the underlying voice models.
The system takes in visual and audio information and responds with speech and video from a dynamic persona. Google highlights lip-sync, expressions and natural turn-taking. For a customer-service or guided-support team, that could make an interaction feel less like a voice menu and more like a conversation. Whether it actually improves completion rates remains a question for deployment testing.
The distinction matters because a convincing face can raise users’ expectations. An avatar might look attentive while a tool call is still running, or appear confident when its answer is uncertain. Organisations should design for clear status and truthful responses, not assume a fluent presentation is evidence that the underlying task has been completed.
Conversation continues while tools work
Google says Live Avatar supports asynchronous tool execution. It can fetch information or trigger a connected action in the background while the dialogue stays active. The company demonstrates a hotel check-in scenario, where the agent continues talking during a background call. That is a meaningful design change from assistants that freeze a conversation whenever they consult another system.
A service team should separate conversation continuity from action authority. Looking up a reservation, changing a room and charging a card have different consequences. The avatar can explain what it is checking, but the application still needs identity verification, bounded tool permissions and confirmation for consequential changes. Those controls should remain effective even when the exchange feels informal.
Latency needs to be measured end to end. A generated video stream, speech recognition, reasoning, network calls and backend systems all contribute to the wait a customer experiences. A useful pilot would measure interruptions, abandoned sessions and successful resolution alongside visual smoothness. Maintaining eye contact is valuable only if the assistant’s factual and transactional work is reliable.
Language reach and branding
Google says the feature can transition across 97 languages, adapting lip-sync and expressions without visual drift. This could help a service organisation support multilingual conversations through one interface. It does not establish that every domain term, accent or local procedure will be handled correctly. Australian teams should test the languages their customers actually use, including code-switching and specialised vocabulary.
The launch includes a library of preset avatars. Google also describes custom avatars generated from a high-quality reference image to preserve a character’s appearance and brand styling, but says that creation is currently available only through enterprise allowlisting. A listed feature should therefore not be presented to every customer as immediately accessible. Eligibility and configuration need direct confirmation.
Custom characters create a governance question beyond visual polish. A company should hold the rights to a reference image and obtain permission from any recognisable person. It should document how the avatar identifies itself as AI, what it may say on the organisation’s behalf and how to retire an identity when a campaign or employee role changes.
Trust is part of the interface
Google says generated audio and video carry SynthID watermarks. The mark is intended to support detection of AI-generated output and reduce misattribution. It is a safeguard, not a substitute for telling people they are speaking with an AI system. An enterprise should disclose that fact in the interface and provide a straightforward handoff when a human conversation is needed.
A practical evaluation should include cases where the camera cannot see a relevant detail, where a user corrects the agent, and where a background system returns an error. The avatar should not mask uncertainty with reassuring gestures. Reviewers can compare a voice-only version with the visual version to establish whether video adds genuine accessibility or service value rather than novelty.
The announcement expands Gemini’s live interface from spoken assistance to an embodied enterprise agent. Its commercial importance will depend on how reliably organisations can pair the new presence with accurate answers, consent, multilingual quality and safe actions. The right launch measure is a completed, understood customer task, not simply a technically impressive character on screen.
The source does not provide a universal per-minute price, a complete country list or a public date for self-service custom avatar creation. Buyers should confirm those points in their own Gemini Enterprise agreement and documentation before promising a launch date. A small trial with realistic user consent is more informative than assuming a research demonstration maps directly to production entitlement.
Accessibility deserves its own test plan. Some users may prefer a visual character, while others may need captions, a text transcript or a simple voice-only option. The interface should let a customer interrupt, repeat an answer and inspect the result of any background action. These choices help make the additional visual layer useful rather than obligatory.