Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new AI models designed to make voice-based AI significantly more capable.
Instead of functioning like a traditional voice assistant that simply listens to a question and returns an answer, these new Gemini models are designed to maintain a natural conversation while reasoning, understanding visual information, calling tools, and completing tasks in the background.
This represents an important shift in conversational AI: from voice assistants toward real-time AI agents.
What Is Gemini 3.8 Live?
Gemini 3.8 Live is Google’s latest native speech-to-speech model designed for scalable, real-time conversational applications.
Google is launching two versions:
Gemini 3.8 Live focuses on speed, conversational intelligence and cost-efficient deployment at scale.
Gemini 3.8 Live Extended Thinking is intended for more complex tasks requiring deeper reasoning and multi-step decision-making.
The key difference is that Extended Thinking can perform more substantial reasoning while still maintaining an ongoing conversation with the user.
That capability could be particularly important for enterprise AI agents handling workflows rather than simple questions.
AI Can Now Work While You Continue Talking
One of the most interesting improvements is asynchronous function calling.
Traditional voice AI interactions can feel sequential:
User speaks → AI processes → tool runs → user waits → AI responds.
Gemini 3.8 Live can instead execute tools and API calls in the background while continuing to interact with the user.
Imagine saying:
“Check my order, confirm whether it has shipped, and reschedule the delivery for Friday.”
The AI could acknowledge the request, continue the conversation and execute the required backend operations without forcing the interaction to stop while each system responds.
For developers building AI agents, this architecture could make voice applications feel substantially more natural.
Gemini Can See While It Talks
Gemini 3.8 Live also supports live visual context.
This means an application can combine what a user is saying with what the AI can currently see.
For example, a technician could point a camera at a machine and ask:
“Which component should I inspect first?”
A Gemini-powered assistant could potentially analyze the live visual input while discussing troubleshooting steps with the technician.
Google demonstrated this capability with scenarios including employee onboarding and even playing chess using visual context while maintaining a live conversation.
Potential applications include:
- Technical field support
- Equipment troubleshooting
- Remote inspections
- Employee training
- Healthcare administration workflows
- Customer service
- Guided product installation
- Accessibility applications
The combination of voice + vision + reasoning + tools is what makes this release particularly significant.
Support for More Than 97 Languages
Gemini 3.8 Live can automatically detect and transition between 97 supported languages during a conversation.
For organizations operating internationally, multilingual conversations could therefore happen without requiring users to manually change language settings.
Google also highlights improved handling of alphanumeric information such as confirmation codes, claim numbers and other technical data — an important requirement for enterprise voice agents.
Extended Thinking Brings Deeper Reasoning to Voice
Gemini 3.8 Live Extended Thinking addresses one of the biggest limitations of traditional voice assistants: complicated reasoning.
Instead of requiring the system to immediately generate an answer, Extended Thinking can perform deeper multi-step reasoning while maintaining the conversation.
For example, an enterprise assistant might receive a request such as:
“Review these three customer issues, check our support policies, determine which cases qualify for escalation, and prepare the next steps.”
That task requires more than speech recognition.
The AI needs to:
Understand the request → retrieve information → reason about policies → potentially call enterprise systems → make decisions → explain the result.
Google says Extended Thinking is specifically designed for these more complex workflows.
Voice Agents Are Becoming Enterprise Agents
This may ultimately be the bigger story behind Gemini 3.8 Live.
Voice AI is moving beyond applications such as:
“What’s the weather?”
toward workflows such as:
“Check the customer’s account, verify their eligibility, update the case and schedule a follow-up appointment.”
The difference is significant.
The first interaction requires an AI assistant.
The second requires an AI agent capable of reasoning and interacting with enterprise systems.
Gemini’s ability to execute tools while continuing a conversation makes these kinds of applications considerably more practical.
Potential Enterprise Use Cases
Organizations could eventually use this type of technology across several areas.
Customer Service
A voice agent could identify the customer, retrieve account information, answer questions and execute actions without transferring between multiple systems.
IT Help Desk
Employees could describe an issue verbally while sharing their screen or camera. The AI could guide troubleshooting while simultaneously checking documentation or creating service tickets.
Government Citizen Services
Residents could potentially interact naturally with government services using voice while the AI retrieves policies, forms, case information and service availability.
Healthcare Administration
Voice agents could assist with appointment scheduling, administrative questions and navigation of approved information sources.
Field Service
Technicians could show equipment to the AI while receiving step-by-step troubleshooting assistance.
Employee Onboarding
New employees could ask questions while navigating applications, documents or internal processes.
These scenarios become significantly more powerful when conversational AI can understand, reason, see and act simultaneously.
How Well Does Gemini 3.8 Live Perform?
Google reports strong benchmark results for the new models.
Gemini 3.8 Live Extended Thinking achieved a score of 82.6 on Artificial Analysis’ Speech-to-Speech Quality Index and ranked first overall at the time of Google’s announcement.
Google also reports scores of:
- 68.6% on τ-Voice
- 35.1% on Sierra’s τ-Voice Banking benchmark
- 97.7% on Big Bench Audio
These are benchmarks rather than guarantees of production performance, so organizations should still evaluate the models against their own workflows, latency requirements, security controls and accuracy expectations.
What Does Gemini 3.8 Live Cost?
For developers using the Gemini Live API, Google lists pricing of approximately:
Audio input: $0.005 per minute
Audio output: $0.018 per minute
Google positions these models as an alternative to more complex voice architectures that separately combine speech recognition, an LLM and text-to-speech systems.
Actual application costs will still depend on architecture, conversation length, tool usage, infrastructure and other services involved.
Where Is Gemini 3.8 Live Available?
Google began rolling out Gemini 3.8 Live on September 15, 2026.
For developers, the models are available through:
- Gemini API
- Google AI Studio
- Gemini Live API
Gemini 3.8 Live is also entering private preview in Gemini Enterprise, with planned availability for Gemini Enterprise for Customer Experience.
Gemini 3.8 Live Extended Thinking is rolling out through Gemini Live and selected Google Workspace experiences, including Gmail, Docs and Keep depending on subscription and availability.
Google is also working with platforms including LiveKit, Pipecat, Agora, Fishjam, Vercel and Vision Agents to support developers building real-time voice applications.
Google Is Adding SynthID to AI-Generated Audio
As realistic AI-generated speech becomes more common, identifying synthetic audio becomes increasingly important.
Google says audio generated by its AI products includes SynthID watermarking.
SynthID embeds an imperceptible watermark into generated audio so that AI-generated content can later be identified.
This type of provenance technology is likely to become increasingly relevant as synthetic voices become harder to distinguish from human speech.
Why Gemini 3.8 Live Matters
The most important part of this announcement isn’t simply improved voice quality.
It is the convergence of multiple AI capabilities into one interaction:
Voice + Vision + Reasoning + Tools + Real-Time Execution
That combination moves AI closer to becoming an interface for software itself.
Instead of opening an application, locating a function, entering information and navigating multiple screens, users may increasingly be able to describe the outcome they want.
The AI agent then determines how to complete the underlying workflow.
That transition could fundamentally change how people interact with enterprise applications over the next several years.
The Bigger Shift: From Chatbots to Real-Time AI Agents
The first generation of enterprise generative AI largely focused on chatbots.
The next generation is increasingly focused on agents.
Gemini 3.8 Live shows what that evolution can look like when voice becomes the interface.
An AI agent can potentially:
Listen → See → Understand → Reason → Use tools → Take action → Continue the conversation.
The important question for organizations is therefore shifting from:
“Can AI answer this question?”
to:
“Which business processes can AI safely complete?”
That requires more than a capable model. Organizations will also need strong identity management, permissions, audit trails, grounding, security controls, human oversight and AI governance.
But Gemini 3.8 Live demonstrates that the underlying technology needed to build much more capable conversational agents is advancing quickly.
Final Thoughts
Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking represent another step toward AI systems that can interact with people more naturally while simultaneously performing real work.
For developers, this creates opportunities to build a new class of voice-first applications.
For enterprises, it opens possibilities across customer service, employee support, field operations, government services and other workflows where conversational interaction could replace traditional interfaces.
The future of enterprise AI may not simply be another chatbot inside an application.
Increasingly, the conversation itself may become the application.
Original Source: Google — Introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking Read the original Google announcement
Published by Google: September 15, 2026
