// AI TOOLS · GEMINI 3.8 LIVE

Gemini 3.8 Live

Google's real-time voice models — a cheap, fast conversational tier plus a reasoning-heavy Extended Thinking variant that thinks and speaks at once, with visual grounding and 97-language switching.

Last verified · 2026-09-15 · by Moe Ameen

What Gemini 3.8 Live is

Gemini 3.8 Live is a Google live audio model announced on September 15, 2026, built for real-time spoken interaction instead of typed prompts. It shipped in two variants: Gemini 3.8 Live, the scale-and-cost tier tuned for fluid dialogue and "visual grounding," and Gemini 3.8 Live Extended Thinking, a higher-intelligence tier for multi-step reasoning. Both are designed for voice agents that feel like a conversation — they can execute tool and API calls in the background while continuing to talk, so nothing stalls while a lookup runs.

The standout capabilities for a creator are the multimodal and multilingual sides. Visual grounding means the model can take a live camera or screen feed and reason about what it sees in near real time — point at a product, a rough cut, or a whiteboard and talk it through. And the models automatically detect and transition between 97 supported languages mid-conversation, so a bilingual creator can switch mid-sentence. Extended Thinking adds a "reasons and speaks simultaneously" behavior — early verbal cues like "Let me check that…" and spoken progress narration while a background task runs.

Google made a price-and-performance case at launch. It reports Extended Thinking topping Artificial Analysis' Speech-to-Speech Quality Index at 82.6, ahead of rival live models, with strong agentic scores on the τ-Voice benchmark. Third-party coverage put Gemini 3.8 Live around $0.84 per hour of input audio and Extended Thinking near $3.50 per hour — below the competing voice models it was measured against. Treat the benchmarks as vendor-reported until independently tested, and confirm live rates in the Gemini API docs. All generated audio carries SynthID watermarking. Access: developers via the Gemini API and Google AI Studio, enterprises via a private preview in Gemini Enterprise, and general users via Search Live, Gemini Live, and Extended Thinking in Google Workspace (Docs, Gmail, Keep) for Pro and Ultra subscribers.

Honest framing: Gemini 3.8 Live is a conversational interface, not a content-production tool. It is excellent for talking through ideas, hands-free research, and multilingual dictation — but it outputs a spoken exchange, not on-brand posts, captioned video, carousels, or a publishing schedule.

What you can make with it

  • Voice-dictated scripts, hooks, and outlines you capture as text — in any of 97 languages, switching mid-conversation
  • A hands-free brainstorm where you talk through an angle out loud and get real-time pushback
  • Live visual walkthroughs — point a camera at a product or a rough edit and narrate the pitch while the model reasons about what it sees
  • Spoken research runs where the model runs background tool calls and narrates progress (Extended Thinking)
  • Multilingual talking points for the same idea, drafted by voice for different audiences
  • A spoken thinking partner for planning a content week while your hands are busy filming or editing

How Kompozy turns Gemini 3.8 Live output into content

Gemini 3.8 Live is where a script gets spoken; it is not where that script becomes a face on camera or a post on nine feeds. Its genuine edge for creators is capture: dictate a hook in any of 97 languages, hold your phone up to a product and talk through the angle with visual grounding, let Extended Thinking narrate its reasoning back. What you get is a clean transcript in the language you spoke. [Kompozy](/) is the engine that turns that transcript into finished, published content — and it is a generation engine, not a scheduler bolted onto a voice app, so it makes the formats Gemini 3.8 Live can't.

The highest-leverage handoff is video. Paste your dictated script into Kompozy and it generates [Persona Shorts](/glossary/persona-shorts) and HeyGen avatar video that put a face-locked, recurring on-camera identity in front of that exact script — so instead of a disembodied voice you get a consistent presenter your audience recognizes week over week, with native TTS and auto-captions. From the same source, one [Persona Brief](/glossary/persona-brief) governs voice and banned phrases while Kompozy fans the idea into brand-exact Carousel Posts and Persona Tweets via [HyperFrames](/glossary/hyperframes), Photo Posts, Quote Graphics, a Blog Article, and an Email Newsletter — then schedules and publishes the whole set across the eight social platforms plus blog and email with [Autopilot](/glossary/autopilot), behind a per-post review gate. Speak the script into Gemini 3.8 Live; put a face on it and ship it everywhere with Kompozy.

  1. Open Gemini 3.8 Live and dictate your script or talk through the angle — use visual grounding to react to a product on camera, or switch language mid-conversation for a specific audience.
  2. Capture the spoken output as text — the transcript, an outline, or the key talking points.
  3. Paste it into Kompozy Quick Ingest as the source for a new content unit.
  4. Set your Persona Brief once, then generate the formats — a face-locked Persona Short reading the script, brand-exact carousels and quote graphics, text posts, a blog, and a newsletter.
  5. Review each piece behind the per-post gate, then schedule and publish across the eight social platforms plus blog and email with Autopilot.

Frequently asked questions

What is Gemini 3.8 Live?

Gemini 3.8 Live is a Google live audio model announced September 15, 2026 for real-time spoken interaction. It offers fluid dialogue, visual grounding (reasoning about a live camera or screen feed in near real time), and background tool calls, and it can switch between 97 languages mid-conversation. A higher-reasoning variant, 3.8 Live Extended Thinking, ships alongside it.

What is the difference between 3.8 Live and 3.8 Live Extended Thinking?

3.8 Live is the cost-efficient, scale tier for fluid dialogue and visual grounding. Extended Thinking is the higher-intelligence tier for multi-step reasoning — Google says it reasons and speaks at the same time, gives early cues like "Let me check that…," and narrates progress out loud. Extended Thinking is reported to cost more per hour of input audio.

Can Gemini 3.8 Live generate video or social posts?

No. It is a conversational voice interface — it talks, reasons, sees, and runs tools, but it generates no avatar video, carousels, images, blogs, or scheduled posts. To turn a dictated script into a face-locked persona video and finished posts across platforms, you use a generation engine like Kompozy, which then also schedules and publishes.

How much does Gemini 3.8 Live cost?

Third-party coverage at launch put Gemini 3.8 Live at roughly $0.84 per hour of input audio and Extended Thinking around $3.50 per hour, below the competing live models it was benchmarked against. Google didn't headline a price in its announcement, so confirm current rates in the Gemini API pricing docs.

Related tools

  • GPT-LiveOpenAI's new full-duplex voice models for ChatGPT — they listen and speak at the same time, and can hand a question off for a web search mid-conversation.
  • Gemini 3.8 FlashGoogle's fast, low-cost workhorse model that 'works harder' on coding, agents, and analysis — launched September 2, 2026 alongside a defense-focused Cyber variant.
  • Gemini Omni FlashGoogle's conversational video model — generate a clip, then refine it by chatting instead of re-prompting.
  • HeyGenAI avatar video platform that turns a text script into a talking-head video — in 175+ languages.
  • Gemini AppGoogle's consumer AI assistant — chat, voice conversations with live camera and screen sharing, image generation, and deep research on Android, iOS, web, and desktop. It crossed one billion monthly active users in August 2026.

← All AI tools · Get started →