Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, as its newest live dialogue models for real-time voice applications. The important change is not only better speech quality. Google is combining native audio-to-audio interaction with multimodal input, background tool execution, asynchronous function calling and a separate model variant for deeper reasoning. For developers, the practical question is how the two models differ, what the Live API actually supports, where the limits are, and what changes when you move from the older Gemini 3.1 Flash Live model.
What Google launched on September 15, 2026
Google launched two related models. gemini-3.8-live is the standard low-latency option for most real-time voice agent experiences. gemini-3.8-live-extended-thinking is the higher-reasoning option for complex, multi-step work during a live conversation.
Both are available through the Gemini API and Google AI Studio. Google also says Gemini 3.8 Live is rolling into Search Live, while the Extended Thinking model is being used in selected Gemini and Workspace experiences. The two models share the same broad Live API interface, but they are meant for different latency and reasoning needs.
This matters because a voice agent does not behave like a normal text request. The application has to keep a live session open, process streaming media, preserve conversation state, handle interruptions and deal with external actions while the user is still speaking. The 3.8 release is built around that problem rather than simply adding speech output to a text model.
What Gemini 3.8 Live can process and return
The current model documentation lists text, images, audio and video as supported input types for gemini-3.8-live. Its outputs are text and audio. The listed input limit is 131,072 tokens and the output limit is 65,536 tokens.
The same documentation lists audio generation, Live API support, search grounding and function calling. It does not list code execution, file search, image generation, URL context or Google Maps grounding for this model. That distinction is useful when planning an agent: the Live model can call functions that your application exposes, but it is not a general-purpose tool box where every Gemini API capability is automatically available.
Google describes the standard model as the default choice for low-latency voice agents and real-time dialogue without reasoning-induced delays. It also supports interleaved reasoning, which means the system can use reasoning while continuing a live interaction instead of turning the conversation into a stop-and-wait text exchange.
How the Extended Thinking model is different
gemini-3.8-live-extended-thinking is aimed at situations where the agent needs more background reasoning. Google describes it as a high-reasoning audio-to-audio model that can process reasoning and tool calls while streaming continuous audio responses.
The model page lists the same 131,072-token input limit and 65,536-token output limit. It supports function calling, but the documented mode is asynchronous. Search grounding is supported, while code execution, file search, Google Maps grounding, image generation, structured outputs and URL context are not listed as supported capabilities.
The practical difference is therefore not simply “better voice.” The choice is tied to task complexity. A short customer question may favor the standard Live model. A live booking flow, troubleshooting session or multi-step task can justify Extended Thinking when the added reasoning is more useful than the extra model work.
What the Live API actually looks like
The Gemini Live API uses WebSockets for persistent, bidirectional communication. Google’s SDK examples show a client opening a live session, selecting a model such as gemini-3.8-live, setting response modalities and then streaming audio or other input into the session.
This architecture is different from a typical request-response API. Instead of sending one message and waiting for one finished answer, the application maintains a session and continuously exchanges data with the model. That is why session state, reconnection and context management are part of the API design rather than optional extras.
Google’s session management documentation includes session resumption. When enabled, the service can issue a resumption token that the client can use after a WebSocket connection is reset. Those resumption tokens are valid for two hours after the last session termination. For a voice application, this is important because a dropped connection should not automatically mean losing the entire conversation state.
Multimodal input is part of the live session
Gemini 3.8 Live is not limited to microphone input. Google’s model page lists images and video alongside audio and text. The launch post also describes near-real-time visual processing, such as an agent using visual context during troubleshooting, onboarding and interactive tasks.
This changes the design of a voice product. A user can talk while the application sends a camera frame, document image or video context into the same live interaction. The model can then respond with audio while using the extra visual signal.
For developers, the important question is not whether the model accepts multiple modalities. It is how often the application should send them. Streaming every frame of a camera feed can increase cost and processing load. A better implementation can send only the visual updates that affect the current task, while keeping the conversation state stable.
Asynchronous function calling changes agent design
One of the more important implementation changes in 3.8 is asynchronous function calling. Google says asynchronous execution using NON_BLOCKING is now the default for gemini-3.8-live, while the older blocking behavior can be selected explicitly. Extended Thinking also uses asynchronous function calling.
In a real voice agent, this allows the conversation to continue while an external action is running. The agent might acknowledge a request, call a booking system or internal API, and keep the dialogue alive instead of forcing the user to wait silently for the tool response.
There is an important boundary here: the Live API documentation says automatic tool response handling is not provided in the same way as the standard generate-content flow. The client application must handle returned tool calls and send the function response back. In other words, your application still owns the orchestration layer.
That means the model should not be treated as the entire agent. The application remains responsible for authentication, authorization, tool validation, retries, timeouts, business rules and the final decision about whether a tool action is allowed.
Search grounding is supported, but not every tool is
The current 3.8 model documentation lists Search grounding as supported for both Live models. At the same time, several other Gemini capabilities are explicitly marked unsupported, including code execution, file search, URL context and Google Maps grounding.
This is useful for architecture decisions. A developer can design a voice assistant that answers with current web information and can call application-defined functions, but should not assume it can directly use every tool available in non-Live Gemini models.
The safest approach is to map every external action to an explicit tool declaration and test the exact model-tool combination used in production. Do not build around a capability because it exists somewhere else in the Gemini API.
Pricing: audio is billed differently from text
Google’s current Gemini API pricing page lists a free tier and paid pricing for the 3.8 Live models. On the paid tier, audio input is listed at $3.00 per million tokens or $0.005 per minute. Audio output is $12.00 per million tokens or $0.018 per minute. Text input is $0.75 per million tokens and text output is $4.50 per million tokens.
The token numbers alone do not tell the whole cost story. Google’s Live API best-practices documentation explains that the service bills the accumulated session context on each turn. As a conversation grows, earlier context can be re-processed and billed again. Long sessions can therefore become more expensive even if the user is not sending much new information per turn.
Google recommends context window compression for long sessions. The example in its documentation uses a compression trigger of 25,000 tokens and a sliding window of 8,000 tokens. The exact settings depend on the application, but the underlying lesson is clear: persistent voice agents need an explicit context-management strategy.
Proactive audio can affect both behavior and billing
Google documents proactive audio as a special behavior for the 3.8 Live models. Proactive audio is permanently enabled for gemini-3.8-live and gemini-3.8-live-extended-thinking.
Under Google’s billing guidance, proactive audio means input tokens can be charged while the Live API is continuously listening, while output tokens are charged when the model responds. This makes always-on listening a product decision as well as a technical one.
For a hands-free assistant, that behavior may be expected. For an application with long idle periods, developers should measure the cost of keeping a session alive rather than estimating the bill from response time alone.
What happens when you migrate from Gemini 3.1 Flash Live
Google’s 3.8 Live documentation includes a direct migration path from gemini-3.1-flash-live-preview. The model string changes to gemini-3.8-live.
There is also a configuration change. Google’s documentation says thinking_level and thinking_config are not supported for the standard 3.8 Live model, so older setup code that sends those fields should be removed.
Function calling also changes. Asynchronous execution using NON_BLOCKING is now the default, while BLOCKING can be explicitly selected when backward-compatible synchronous behavior is needed.
Google’s deprecation page lists Gemini 3.1 Flash Live as a legacy preview model and recommends updating to Gemini 3.8 Live. It does not list a shutdown date for 3.1 Flash Live, so migration is a recommendation rather than a forced immediate shutdown.
Security should be handled outside the model
Google’s Live API overview recommends ephemeral tokens for client-to-server production deployments rather than putting a standard API key in a browser or client application. The ephemeral-token documentation shows that a temporary token can be restricted to a specific Live API model and configuration.
That design is especially relevant for voice agents because the session can call tools and receive continuous user input. The model can decide to request an action, but your server should still validate whether that action is allowed, for which user, with which arguments and under which account or permission.
The Live API also exposes session-resumption features, which means a robust application should plan for reconnects rather than assuming a WebSocket will stay open forever. Logging, rate limits, tool timeouts and audit records should be implemented at the application layer.
How the two 3.8 Live models compare
| Area | Gemini 3.8 Live | Gemini 3.8 Live Extended Thinking |
|---|---|---|
| Primary use | Low-latency live dialogue and most voice agents | Complex, multi-step real-time tasks |
| Model ID | gemini-3.8-live | gemini-3.8-live-extended-thinking |
| Inputs | Text, images, audio, video | Text, images, audio, video |
| Outputs | Text and audio | Text and audio |
| Input limit | 131,072 tokens | 131,072 tokens |
| Output limit | 65,536 tokens | 65,536 tokens |
| Function calling | Supported; asynchronous default | Supported; asynchronous |
| Search grounding | Supported | Supported |
| Code execution | Not supported | Not supported |
| File search | Not supported | Not supported |
| Structured outputs | Not supported | Not supported |
Where Gemini 3.8 Live fits compared with other voice architectures
The launch also highlights an architectural difference that matters beyond model quality. An independent analysis from The New Stack describes Google’s approach as keeping reasoning inside the live voice model, while OpenAI’s GPT-Live-1 approach can separate real-time conversation from a backend reasoning model.
That distinction does not prove that one architecture is better for every workload. It changes where complexity lives. With an integrated live reasoning model, more of the conversational state and task behavior sits close to the voice session. With a split design, more orchestration can live in your application. The choice affects latency, state management, tool routing, cost accounting and failure handling.
For developers evaluating both approaches, test the full application rather than comparing voice demos. Record task completion, interruption handling, time to first audio, tool-call latency, failed actions, reconnect behavior and total cost per completed task.
What the September 15 launch actually changes for developers
Three changes stand out when the launch announcement is read together with the current API documentation.
First, live multimodal interaction is now a stable model path rather than only a narrow preview experience. The model page lists gemini-3.8-live as stable, and the deprecation page lists the release date as September 15, 2026 with no shutdown date announced.
Second, asynchronous tool use is part of the main agent design. A voice application can continue the dialogue while an external function is running, but the client still has to manage the tool response.
Third, the economics are tied to persistent session context. Developers need to treat context compression, idle listening and token growth as engineering concerns, not just billing details.
What remains unconfirmed or easy to misunderstand
Google’s public documentation explains the public 3.8 Live models and Live API capabilities. It does not establish that every internal Google voice or agent deployment uses the same configuration. Public API support should not be treated as evidence about private internal systems.
Google’s launch article also reports benchmark results from Artificial Analysis, ServiceNow EVA-Bench and other evaluations. Those measurements are useful context, but benchmark performance does not tell a developer how the model will behave on a particular production workflow. The right validation target is the actual workload, including interruptions, tool failures, noisy audio and long sessions.
How to test Gemini 3.8 Live before production
Start with a fixed test set of real conversations. Include normal requests, interrupted requests, ambiguous speech, visual context, unavailable tools, slow tools, rejected actions and reconnects. Test both the standard Live model and Extended Thinking on the same scenarios where both are viable.
Measure task completion, time to first response, total turn latency, audio input and output usage, context growth, tool-call errors, recovery after disconnects and human corrections. Also test a long session with context compression enabled and compare the cost with an uncompressed session.
For browser clients, implement ephemeral-token issuance on the server and keep the long-lived API credential out of client code. For tools that can create side effects, use explicit allowlists, argument validation, authorization checks and application-level logging.
Related ToolBoxKart guides
For the non-Live Gemini model family, see Gemini 3.8 Flash API features and pricing. For live voice architecture, see GPT-Live-1 Voice API. For agent testing, read AI Agent Independent Evaluation. For permissions, see How to Audit AI Agent Permissions. For broader agent architecture, see AI Agent Architect. For another Gemini workflow, read the Gemini Windows app guide.
Frequently asked questions
What is the Gemini 3.8 Live model ID?
The standard model ID is gemini-3.8-live. The higher-reasoning variant is gemini-3.8-live-extended-thinking.
Does Gemini 3.8 Live support video input?
Yes. Google's current model documentation lists video alongside text, images and audio as supported input types.
Does Gemini 3.8 Live support function calling?
Yes. Google documents function calling for the Live models and asynchronous execution for live tool use.
Sources
- Google — Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking
- Google AI for Developers — Live API
- Google AI for Developers — Gemini 3.8 Live model documentation
- Google AI for Developers — Live API guide and best practices
- Google AI for Developers — Ephemeral tokens
- Google AI for Developers — Deprecations
- The Register — Google Gemini 3.8 Live coverage
- The New Stack — Gemini 3.8 Live architecture analysis