Google shipped two live audio models — a cheap, fast conversational tier and a reasoning-heavy "Extended Thinking" variant that thinks and speaks at the same time — aimed at natural voice agents rather than turn-based assistants.
2026-09-15 · by Moe Ameen
Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026 — two live audio models built for real-time, spoken interaction rather than typed prompts. The pitch is voice agents that feel like a conversation: the models listen and respond in the flow of dialogue, and can run tool and API calls in the background while they keep talking, so there's no dead air while something loads.
The two variants split by job. Gemini 3.8 Live is the scale-and-cost tier — tuned for fluid dialogue and "visual grounding," meaning it can take in a live camera or screen feed and reason about what it sees in near real time. Gemini 3.8 Live Extended Thinking is the higher-intelligence tier for multi-step tasks: Google says it reasons and speaks simultaneously, offering early verbal cues like "Let me check that…" and narrating progress out loud while a background task runs. Both models can automatically detect and switch between 97 languages mid-conversation.
Google leaned on benchmarks to make a price-and-performance case. It reports Extended Thinking topping Artificial Analysis' Speech-to-Speech Quality Index at 82.6, ahead of rival voice models, with strong agentic scores on the τ-Voice benchmark. On price, third-party coverage put Gemini 3.8 Live at roughly $0.84 per hour of input audio and Extended Thinking around $3.50 per hour — materially below the competing live models it was measured against. Treat the benchmark figures as vendor-reported until independent testing lands, and confirm live pricing in the Gemini API docs. All AI-generated audio carries SynthID watermarking.
Availability rolled out the same day: developers get both models through the Gemini API and Google AI Studio; enterprises via a private preview in Gemini Enterprise; and general users through Search Live, Gemini Live, and Extended Thinking inside Google Workspace apps (Docs, Gmail, Keep) for Pro and Ultra subscribers. The honest creator framing: this is a conversational interface, not a content-production tool — it holds a great spoken exchange, but it exports no captioned video, carousel, blog, or scheduled post.
The reason a cheaper, more natural voice model matters to a creator isn't the conversation — it's the capture. Gemini 3.8 Live is the fastest way yet to get an idea out of your head: talk through a hook while you're driving, hold your phone up to a product and narrate the angle, let it switch to Spanish mid-sentence for a bilingual audience. What you walk away with is a transcript and a sharper take. That's exactly where [Kompozy](/) begins. Paste the transcript or the tightened angle into Quick Ingest and one spoken brainstorm fans out into a full week: a Blog Article, brand-exact Carousel Posts and Quote Graphics via [HyperFrames](/glossary/hyperframes), native Text Posts, and an Email Newsletter — every piece rewritten to your voice by the [Persona Brief](/glossary/persona-brief) and banned-word filters, so a rambling voice note reads as your brand rather than a raw dictation.
Then Kompozy does the two things a voice model structurally can't. It generates the video: [Persona Shorts](/glossary/persona-shorts) and HeyGen avatar clips that put a face-locked, recurring on-camera identity in front of the script you just dictated, plus [Clipped Shorts](/glossary/clipped-short) if you already have footage. And it publishes: the whole set is scheduled and fanned across the eight social platforms plus blog and email through [Autopilot](/glossary/autopilot), behind a per-post review gate. Say the idea into Gemini 3.8 Live; ship it everywhere with Kompozy.
Gemini 3.8 Live is a Google live audio model announced September 15, 2026 for real-time spoken interaction — fluid dialogue, visual grounding (reasoning about a live camera or screen feed in near real time), and the ability to run tool and API calls in the background while it keeps talking. It ships alongside a higher-reasoning variant, 3.8 Live Extended Thinking.
Extended Thinking is the higher-intelligence tier for multi-step tasks. Google says it reasons and speaks at the same time, gives early verbal cues like "Let me check that…," and narrates progress while a background task runs. It is reported to cost more per hour of input audio than the base 3.8 Live model, in exchange for stronger reasoning and agentic performance.
No. It is a conversational voice interface — it talks, reasons, sees, and can run tools in the background, but it produces no captioned video, carousel, blog, newsletter, or scheduled post. To turn a spoken idea or transcript into finished, on-brand posts across platforms, you run it through a content engine like Kompozy.
Third-party coverage at launch put Gemini 3.8 Live at roughly $0.84 per hour of input audio and 3.8 Live Extended Thinking around $3.50 per hour — below the competing live voice models it was benchmarked against. Google didn't headline a price in its announcement, so confirm current rates in the Gemini API pricing docs.