Google's real-time voice models — a cheap, fast conversational tier plus a reasoning-heavy Extended Thinking variant that thinks and speaks at once, with visual grounding and 97-language switching.
Last verified · 2026-09-15 · by Moe Ameen
Gemini 3.8 Live is a Google live audio model announced on September 15, 2026, built for real-time spoken interaction instead of typed prompts. It shipped in two variants: Gemini 3.8 Live, the scale-and-cost tier tuned for fluid dialogue and "visual grounding," and Gemini 3.8 Live Extended Thinking, a higher-intelligence tier for multi-step reasoning. Both are designed for voice agents that feel like a conversation — they can execute tool and API calls in the background while continuing to talk, so nothing stalls while a lookup runs.
The standout capabilities for a creator are the multimodal and multilingual sides. Visual grounding means the model can take a live camera or screen feed and reason about what it sees in near real time — point at a product, a rough cut, or a whiteboard and talk it through. And the models automatically detect and transition between 97 supported languages mid-conversation, so a bilingual creator can switch mid-sentence. Extended Thinking adds a "reasons and speaks simultaneously" behavior — early verbal cues like "Let me check that…" and spoken progress narration while a background task runs.
Google made a price-and-performance case at launch. It reports Extended Thinking topping Artificial Analysis' Speech-to-Speech Quality Index at 82.6, ahead of rival live models, with strong agentic scores on the τ-Voice benchmark. Third-party coverage put Gemini 3.8 Live around $0.84 per hour of input audio and Extended Thinking near $3.50 per hour — below the competing voice models it was measured against. Treat the benchmarks as vendor-reported until independently tested, and confirm live rates in the Gemini API docs. All generated audio carries SynthID watermarking. Access: developers via the Gemini API and Google AI Studio, enterprises via a private preview in Gemini Enterprise, and general users via Search Live, Gemini Live, and Extended Thinking in Google Workspace (Docs, Gmail, Keep) for Pro and Ultra subscribers.
Honest framing: Gemini 3.8 Live is a conversational interface, not a content-production tool. It is excellent for talking through ideas, hands-free research, and multilingual dictation — but it outputs a spoken exchange, not on-brand posts, captioned video, carousels, or a publishing schedule.
Gemini 3.8 Live is where a script gets spoken; it is not where that script becomes a face on camera or a post on nine feeds. Its genuine edge for creators is capture: dictate a hook in any of 97 languages, hold your phone up to a product and talk through the angle with visual grounding, let Extended Thinking narrate its reasoning back. What you get is a clean transcript in the language you spoke. [Kompozy](/) is the engine that turns that transcript into finished, published content — and it is a generation engine, not a scheduler bolted onto a voice app, so it makes the formats Gemini 3.8 Live can't.
The highest-leverage handoff is video. Paste your dictated script into Kompozy and it generates [Persona Shorts](/glossary/persona-shorts) and HeyGen avatar video that put a face-locked, recurring on-camera identity in front of that exact script — so instead of a disembodied voice you get a consistent presenter your audience recognizes week over week, with native TTS and auto-captions. From the same source, one [Persona Brief](/glossary/persona-brief) governs voice and banned phrases while Kompozy fans the idea into brand-exact Carousel Posts and Persona Tweets via [HyperFrames](/glossary/hyperframes), Photo Posts, Quote Graphics, a Blog Article, and an Email Newsletter — then schedules and publishes the whole set across the eight social platforms plus blog and email with [Autopilot](/glossary/autopilot), behind a per-post review gate. Speak the script into Gemini 3.8 Live; put a face on it and ship it everywhere with Kompozy.
Gemini 3.8 Live is a Google live audio model announced September 15, 2026 for real-time spoken interaction. It offers fluid dialogue, visual grounding (reasoning about a live camera or screen feed in near real time), and background tool calls, and it can switch between 97 languages mid-conversation. A higher-reasoning variant, 3.8 Live Extended Thinking, ships alongside it.
3.8 Live is the cost-efficient, scale tier for fluid dialogue and visual grounding. Extended Thinking is the higher-intelligence tier for multi-step reasoning — Google says it reasons and speaks at the same time, gives early cues like "Let me check that…," and narrates progress out loud. Extended Thinking is reported to cost more per hour of input audio.
No. It is a conversational voice interface — it talks, reasons, sees, and runs tools, but it generates no avatar video, carousels, images, blogs, or scheduled posts. To turn a dictated script into a face-locked persona video and finished posts across platforms, you use a generation engine like Kompozy, which then also schedules and publishes.
Third-party coverage at launch put Gemini 3.8 Live at roughly $0.84 per hour of input audio and Extended Thinking around $3.50 per hour, below the competing live models it was benchmarked against. Google didn't headline a price in its announcement, so confirm current rates in the Gemini API pricing docs.