// AI NEWS · MODEL RELEASE

Google Launches Gemini 3.8 Live and 3.8 Live Extended Thinking, Its Real-Time Voice Models

Google shipped two live audio models — a cheap, fast conversational tier and a reasoning-heavy "Extended Thinking" variant that thinks and speaks at the same time — aimed at natural voice agents rather than turn-based assistants.

2026-09-15 · by Moe Ameen

What happened

Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026 — two live audio models built for real-time, spoken interaction rather than typed prompts. The pitch is voice agents that feel like a conversation: the models listen and respond in the flow of dialogue, and can run tool and API calls in the background while they keep talking, so there's no dead air while something loads.

The two variants split by job. Gemini 3.8 Live is the scale-and-cost tier — tuned for fluid dialogue and "visual grounding," meaning it can take in a live camera or screen feed and reason about what it sees in near real time. Gemini 3.8 Live Extended Thinking is the higher-intelligence tier for multi-step tasks: Google says it reasons and speaks simultaneously, offering early verbal cues like "Let me check that…" and narrating progress out loud while a background task runs. Both models can automatically detect and switch between 97 languages mid-conversation.

Google leaned on benchmarks to make a price-and-performance case. It reports Extended Thinking topping Artificial Analysis' Speech-to-Speech Quality Index at 82.6, ahead of rival voice models, with strong agentic scores on the τ-Voice benchmark. On price, third-party coverage put Gemini 3.8 Live at roughly $0.84 per hour of input audio and Extended Thinking around $3.50 per hour — materially below the competing live models it was measured against. Treat the benchmark figures as vendor-reported until independent testing lands, and confirm live pricing in the Gemini API docs. All AI-generated audio carries SynthID watermarking.

Availability rolled out the same day: developers get both models through the Gemini API and Google AI Studio; enterprises via a private preview in Gemini Enterprise; and general users through Search Live, Gemini Live, and Extended Thinking inside Google Workspace apps (Docs, Gmail, Keep) for Pro and Ultra subscribers. The honest creator framing: this is a conversational interface, not a content-production tool — it holds a great spoken exchange, but it exports no captioned video, carousel, blog, or scheduled post.

Why it matters for creators

  • Real-time voice interaction just got cheap. At a reported ~$0.84/hour of input audio for the base model, voice-first ideation and hands-free research stop being a novelty and become a practical part of a creator's daily workflow.
  • Visual grounding means you can point a camera at a product, a whiteboard, or a rough edit and talk through it live — a faster way to capture an idea than typing a prompt.
  • Mid-conversation language switching across 97 languages makes it a real tool for creators working multilingual audiences, at least at the ideation and script stage.
  • It outputs a conversation, not content. A spoken brainstorm is raw material; it still has to be written per platform, designed, produced as video, captioned, scheduled, and posted before anyone sees it.
  • The benchmarks and pricing are Google- and third-party-reported at launch — judge quality on your own use, and confirm current API rates before you build a workflow around a number.

How to act on this with Kompozy

The reason a cheaper, more natural voice model matters to a creator isn't the conversation — it's the capture. Gemini 3.8 Live is the fastest way yet to get an idea out of your head: talk through a hook while you're driving, hold your phone up to a product and narrate the angle, let it switch to Spanish mid-sentence for a bilingual audience. What you walk away with is a transcript and a sharper take. That's exactly where [Kompozy](/) begins. Paste the transcript or the tightened angle into Quick Ingest and one spoken brainstorm fans out into a full week: a Blog Article, brand-exact Carousel Posts and Quote Graphics via [HyperFrames](/glossary/hyperframes), native Text Posts, and an Email Newsletter — every piece rewritten to your voice by the [Persona Brief](/glossary/persona-brief) and banned-word filters, so a rambling voice note reads as your brand rather than a raw dictation.

Then Kompozy does the two things a voice model structurally can't. It generates the video: [Persona Shorts](/glossary/persona-shorts) and HeyGen avatar clips that put a face-locked, recurring on-camera identity in front of the script you just dictated, plus [Clipped Shorts](/glossary/clipped-short) if you already have footage. And it publishes: the whole set is scheduled and fanned across the eight social platforms plus blog and email through [Autopilot](/glossary/autopilot), behind a per-post review gate. Say the idea into Gemini 3.8 Live; ship it everywhere with Kompozy.

Quick takeaways

  • Google launched Gemini 3.8 Live and 3.8 Live Extended Thinking on September 15, 2026 — live audio models for real-time spoken interaction with background tool calls and visual grounding.
  • 3.8 Live is the cheap, fast conversational tier; Extended Thinking is the reasoning-heavy variant that thinks and speaks simultaneously and narrates progress out loud.
  • Both switch between 97 languages mid-conversation; Google reports Extended Thinking topping Artificial Analysis' Speech-to-Speech index at 82.6; input audio is reported around $0.84/hour (base) and ~$3.50/hour (Extended Thinking).
  • It's a conversational interface, not a content tool — the strategic move is to capture the idea by voice, then generate and publish it with an engine like Kompozy.

Frequently asked questions

What is Gemini 3.8 Live?

Gemini 3.8 Live is a Google live audio model announced September 15, 2026 for real-time spoken interaction — fluid dialogue, visual grounding (reasoning about a live camera or screen feed in near real time), and the ability to run tool and API calls in the background while it keeps talking. It ships alongside a higher-reasoning variant, 3.8 Live Extended Thinking.

How is Gemini 3.8 Live Extended Thinking different?

Extended Thinking is the higher-intelligence tier for multi-step tasks. Google says it reasons and speaks at the same time, gives early verbal cues like "Let me check that…," and narrates progress while a background task runs. It is reported to cost more per hour of input audio than the base 3.8 Live model, in exchange for stronger reasoning and agentic performance.

Can Gemini 3.8 Live create or publish social media content?

No. It is a conversational voice interface — it talks, reasons, sees, and can run tools in the background, but it produces no captioned video, carousel, blog, newsletter, or scheduled post. To turn a spoken idea or transcript into finished, on-brand posts across platforms, you run it through a content engine like Kompozy.

How much does Gemini 3.8 Live cost?

Third-party coverage at launch put Gemini 3.8 Live at roughly $0.84 per hour of input audio and 3.8 Live Extended Thinking around $3.50 per hour — below the competing live voice models it was benchmarked against. Google didn't headline a price in its announcement, so confirm current rates in the Gemini API pricing docs.

Related news

← All AI news · Get started →