Summary
Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, models designed for more natural voice conversations, visual context and background task execution. The company says the Extended Thinking model adds deeper multi-step reasoning while keeping conversations fluid.
Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026, describing them as new models for more natural, capable voice interaction. The models are designed to process speech and visual context in near real time, execute tools in the background and support more complex conversations without forcing the user to stop speaking.
Google says the features are being made available through the Gemini API, Google Workspace and the Gemini app, with Gemini interactions also being expanded across Search.
Two models for different voice workloads
Gemini 3.8 Live is the lower-cost, scale-oriented model. Google describes it as combining conversational intelligence with fluid dialogue and visual grounding, allowing the model to use what it sees as part of a spoken interaction.
The model can process visual inputs in near real time and automatically switch between 97 supported languages during a conversation. It can also run tools and API calls in the background while continuing to talk. For example, it can acknowledge a request, keep the conversation moving and complete the associated task as the operation finishes.
Gemini 3.8 Live Extended Thinking is aimed at more demanding, multi-step work. It can reason and speak simultaneously, using short verbal cues such as “Let me check that…” while it works through a request. Google says the model can narrate progress during background tasks instead of leaving the conversation silent while a complex operation runs.
The company’s demonstrations show Gemini 3.8 Live using visual context during an employee-onboarding interaction and a near-real-time chess game. A separate demonstration shows Gemini 3.8 Live Extended Thinking turning raw sketches and spoken feedback into functional React components.
Reported performance and practical significance
Google reports that Gemini 3.8 Live Extended Thinking achieved a score of 82.6 on Artificial Analysis’ Speech to Speech Quality Index, which the company says placed it first overall. Google also reports scores of 68.6% on the τ-Voice agentic task-completion benchmark and 35.1% on Sierra’s τ-Voice-banking benchmark.
For reasoning, Google says the model scored 97.7% on Big Bench Audio. These results are presented as evidence of performance across spoken interaction, task completion and audio reasoning, rather than as a single measure of voice quality.
Google also says Gemini 3.8 Live ranked second in the Speech Agent Arena and is designed to offer a capable, cost-efficient option for developers and enterprises building voice agents. On ServiceNow’s EVA-Bench, the company says the models balanced conversational quality and accuracy for complex workflows. Google notes that this evaluation was run through the Live API on Gemini Enterprise Agent Platform.
The technical change is the combination of conversation, perception and asynchronous task execution. A conventional voice exchange generally treats speaking and task execution as sequential steps. These models are designed to keep the dialogue active while visual processing, tool use or deeper reasoning continues in the background. That could make voice agents more useful for workflows in which users need both immediate acknowledgement and a completed action.
The benchmark figures are Google-reported results, and the announcement identifies the tests and scores without detailing their full evaluation setup. Product-level access, pricing and rollout timing may also vary across the Gemini app, Workspace, Search and API surfaces.
