Gemini 3.8 Live: Google's Voice AI That Thinks as It Speaks
On Monday, Google shipped two models that make talking to your phone feel like talking to a person. Not a chatbot. Not a voice assistant that interrupts you mid-sentence and makes you repeat yourself. A conversational partner that listens while it works, switches languages mid-sentence, and sees what you are pointing at in real time.
The models are called Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. They are the most advanced voice AI models ever released — and they represent a shift in how AI interfaces with humans. Not through text boxes and keyboards, but through the most natural interface humans have: speech.
What the models do
There are two variants, optimized for different use cases:
Gemini 3.8 Live is built for scale and cost efficiency. It handles most voice interactions: fast, fluid conversations with real-time visual support. It can process what your camera sees, respond to what you are pointing at, and handle interruptions naturally. It executes tool calls and API calls in the background while continuing the conversation — so you do not have to wait while it looks something up.
Gemini 3.8 Live Extended Thinking is built for high-complexity tasks. It reasons and speaks simultaneously. When you ask it something complicated, it uses early verbal cues like “Let me check that…” to acknowledge your prompt naturally, then provides live progress narration as it works through multi-step background tasks. You hear it thinking.
The key technical capabilities:
| Capability | What it means |
|---|---|
| 97 languages | Automatically detects and switches between 97 supported languages mid-conversation |
| Visual grounding | Processes camera input in near real-time; connects verbal responses to specific visual elements |
| Parallel reasoning | Thinks and speaks simultaneously; does not freeze while processing |
| Background execution | Runs tool calls and API calls while continuing the conversation |
| Natural interruptions | Handles mid-sentence interruptions without losing context |
| Live progress narration | Walks you through multi-step tasks as they happen |
The benchmarks
Google did not release these models into a vacuum. They come with benchmark numbers that matter:
| Benchmark | Gemini 3.8 Live Extended Thinking | Competitors |
|---|---|---|
| Speech to Speech Quality Index (Artificial Analysis) | 82.6 (#1 overall) | — |
| τ-Voice (agentic task completion) | 68.6% | — |
| τ-Voice-banking (Sierra) | 35.1% | — |
| Big Bench Audio | 97.7% | — |
| EVA-Bench (ServiceNow) | Pushes Pareto frontier for complex workflows | — |
| Speech Agent Arena | 2nd place user preference | — |
The Speech to Score is the headline number. 82.6 on the Artificial Analysis Speech to Speech Quality Index puts Gemini 3.8 Live Extended Thinking at #1 overall, ahead of every competing voice model.
What this means for developers
For developers, the important detail is the Gemini Live API. It provides the building blocks for production-ready voice agents — not demos, not prototypes, but systems that can handle real conversations at scale.
Developer platforms already supporting Gemini 3.8 Live:
- Agora — real-time media streaming infrastructure
- LiveKit — open-source WebRTC
- Pipecat — voice agent framework
- Vercel — deployment platform
- Vision Agents — visual AI development
The enterprise angle is also significant. Salesforce announced ClaudeForce at Dreamforce this week — a partnership that lets companies use Claude as their AI interface while keeping data in Salesforce’s infrastructure. Google is competing directly with Gemini Enterprise for Customer Experience, which uses Gemini 3.8 Live Extended Thinking as its backbone.
The voice AI landscape in September 2026
Google is not alone in the voice AI space. Here is where things stand:
| Model | Lab | Focus | Status |
|---|---|---|---|
| Gemini 3.8 Live | Voice + visual + reasoning | Released Sep 15 | |
| Gemini 3.8 Live Extended Thinking | Complex voice tasks | Released Sep 15 | |
| GPT-6 Astra | OpenAI | Text + coding + cybersecurity | Released Sep 3 |
| Claude Fable 5.1 | Anthropic | Text + knowledge work | Released Sep 1 |
| Koa | Salesforce + Nvidia | Sales/marketing reasoning | Announced Sep 15 |
The trend is clear: voice AI is no longer a side feature. It is becoming the primary interface for AI interaction. Google, OpenAI, and Anthropic are all investing heavily in real-time conversational AI, and the enterprise market is following.
What this means for users
If you use Google Workspace, Search, or the Gemini app, these models are rolling out to you now:
- Gemini 3.8 Live Extended Thinking is available in the Gemini app for Google AI Pro and Ultra subscribers, and in Workspace for Docs, Gmail, and Keep.
- Gemini 3.8 Live is available in Search Live and rolling out to the Gemini app.
- For developers: Both models are available in the Gemini API and Google AI Studio starting today.
The practical difference you will notice: conversations feel more natural. The model does not freeze while processing. It acknowledges your request and keeps talking while it works. It can see what you are showing it. And it can switch languages mid-sentence without losing context.
The bigger picture
Voice AI is the next interface. Text-based AI changed how we search, write, and code. Voice AI will change how we interact with devices, services, and each other. Google’s release of Gemini 3.8 Live is not just a model update — it is a statement about where the interface is going.
The fact that Google shipped this alongside Gemini 3.8 Live Extended Thinking — a model that reasons while it speaks — suggests that the company sees voice as the primary interface for complex AI tasks, not just simple queries. When you can talk to your AI and it can see what you see, think while it speaks, and act in the background, the text box becomes optional.
This is the beginning of the voice-first era of AI. Google just shipped the best voice model ever made. The question is not whether voice will replace text — it is how fast.
Sources
- Google, “Introducing Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking,” September 15, 2026
- Artificial Analysis, Speech to Speech Quality Index
- ServiceNow, EVA-Bench results
- Sierra, τ-Voice-banking benchmark
- Salesforce, Dreamforce 2026 announcements (ClaudeForce, Koa)
- TechCrunch, “Salesforce and Nvidia’s new reasoning model,” September 15, 2026
Previously on Father of AI
- Why Every AI CEO Is Saying ‘Slow Down’
- The September 2026 AI Model War
- AI Agents Invented Their Own Language
- Father of AI
Frequently asked questions
What is Gemini 3.8 Live?
Gemini 3.8 Live is Google’s most advanced voice AI model, released September 15, 2026. It comes in two variants: Gemini 3.8 Live (optimized for scale and cost efficiency) and Gemini 3.8 Live Extended Thinking (optimized for high-complexity tasks with multi-step reasoning). Both models process visual inputs in near real-time, handle interruptions naturally, switch between 97 supported languages mid-conversation, and execute tool calls in the background while continuing the conversation.
What is the difference between Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking?
Gemini 3.8 Live is built for scale and cost efficiency — it combines conversational intelligence with fluid dialogue and visual grounding. It handles most voice interactions well. Gemini 3.8 Live Extended Thinking is built for high-complexity tasks: it reasons and speaks simultaneously, using early verbal cues like “Let me check that…” to acknowledge prompts naturally, and provides live progress narration as it works through multi-step background tasks. Extended Thinking scored 82.6 on Artificial Analysis’ Speech to Speech Quality Index (#1 overall) and 68.6% on τ-Voice for agentic task completion.
How many languages does Gemini 3.8 Live support?
Gemini 3.8 Live supports 97 languages and can automatically detect and switch between them mid-conversation. This means you can start a conversation in English, switch to Hindi, then switch to Mandarin, and the model will follow without losing context. No other voice AI model supports this many languages with real-time switching.
What can Gemini 3.8 Live see?
Gemini 3.8 Live processes visual inputs in near real-time, enriching conversations with context for more helpful responses. It can see what you are pointing at, read text on a screen, understand diagrams, and ground its responses in what is visible in your camera feed. This is called “visual grounding” — the model connects its verbal responses to specific visual elements in the scene.
Where can I use Gemini 3.8 Live?
Gemini 3.8 Live is rolling out across Google’s ecosystem: the Gemini app (for consumers), Google Workspace (Docs, Gmail, Keep for subscribers), Google Search (Search Live), and the Gemini API and Google AI Studio (for developers). For enterprises, it is available in private preview through Gemini Enterprise and coming soon to Gemini Enterprise for Customer Experience. Developer platforms including Agora, LiveKit, Pipecat, Vercel, and Vision Agents also support it.
How does Gemini 3.8 Live compare to GPT-6 Astra and Claude Fable 5.1?
Gemini 3.8 Live leads in voice-specific benchmarks: #1 on Artificial Analysis’ Speech to Speech Quality Index (82.6), #1 on τ-Voice for agentic task completion (68.6%), and #1 on Big Bench Audio (97.7%). However, these are voice-optimized models — they are not directly comparable to GPT-6 Astra or Claude Fable 5.1 on text-only benchmarks. GPT-6 Astra leads on coding and cybersecurity benchmarks. Claude Fable 5.1 leads on knowledge work and writing. Gemini 3.8 Live leads on voice interaction and multimodal real-time tasks. The right model depends on your use case.