// AI NEWS · MODEL RELEASE

Google Launches Gemini 3.8 Flash TTS and Flash-Lite TTS: Directable AI Voices Across 100+ Languages

Google's most expressive text-to-speech models yet let you design a voice from a description, clone one from 30 seconds of audio, and stage two-speaker dialogue from a single script.

2026-09-23 · by Moe Ameen

What happened

On September 23, 2026, Google introduced two new text-to-speech models — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS — which it calls its most expressive audio-generation models to date. Both turn written scripts into spoken audio across more than 100 languages and dialects (Flash TTS covers the widest set; Flash-Lite slightly fewer), and both are rolling out through the Gemini API and Google AI Studio, with Gemini Enterprise access to follow.

The pitch is control. Beyond a base of 30 studio voices and a library of more than 2,000 ready-to-use voices — including regional variants like Mexican Spanish, Quebec French, and Scots English — you can design a new voice from a plain-text description, or clone one from a 30-second sample (the person being cloned has to record a spoken consent statement first). You can direct delivery line by line for pacing and emotion, stage a two-speaker conversation from a single script, and script non-verbal cues like laughter and sighs. The models handle long-form generation, holding a voice steady across hours of audio, and Google says Flash TTS topped Hume AI's Voice Design benchmark.

The split between the two models is about cost and job. Flash TTS is aimed at creative work — podcasts, audiobooks, game characters — while Flash-Lite TTS targets low-cost speech at scale: dubbing, bulk audio, and voice agents. Promotional pricing through the end of 2026 runs $0.50 per million input tokens plus $9.00 per million tokens of audio output on Flash TTS and $6.00 on Flash-Lite; because Google counts one second of audio as 25 tokens, that works out to roughly $0.81 and $0.54 per hour of generated speech. Every clip carries an inaudible SynthID watermark for AI detection. Prices, language counts, and features change fast on new models — confirm current figures in Google's docs before you build on them.

Why it matters for creators

  • Expressive, directable voice is now cheap enough to narrate at volume — an hour of generated speech for well under a dollar means voiceover stops being the bottleneck for faceless video, audiobooks, and podcasts.
  • Two-speaker dialogue from a single script makes AI podcasts and conversational explainers producible solo, without booking or scheduling a second voice.
  • 100+ languages plus voice cloning let one creator ship the same episode or narration in many markets in the host's own voice — localization without a studio.
  • SynthID watermarking is baked into every clip, so disclosure and detection travel with the audio — relevant as platforms tighten AI-content labeling rules.
  • The raw output is still just an audio file. A voiced podcast or narration isn't a content presence until it's turned into clips, captions, and posts across platforms — the work that comes after the voice.

How to act on this with Kompozy

The takeaway for a creator isn't "switch your voice tool" — it's that the voice layer just got cheap and expressive, which pushes the bottleneck downstream. A polished narration, a two-speaker podcast, or a localized audiobook chapter made with Gemini 3.8 TTS is an audio file, not a channel. Turning it into a week of content is where the work still lives, and that's what [Kompozy](/) is built for. Kompozy doesn't generate the raw voice itself — its own persona and avatar video uses HeyGen's native multi-language TTS — but it is the engine that turns a voiced asset into a published, multi-format stream.

Concretely: pair a Gemini-voiced recording with footage or a persona and Kompozy clips the long piece into captioned vertical shorts, reframes each for TikTok, Reels, and Shorts, then spins the same source into Carousel Posts, a Blog Article, an Email Newsletter, and native Text Posts — all held to one voice by a Persona Brief. [Autopilot](/glossary/autopilot) and a per-post review pipeline schedule and publish the batch across the eight social platforms plus blog and email from a single queue. See [AI voice generation](/glossary/ai-voice-generation) for how the voice layer fits the wider workflow, and the [Gemini 3.8 Live launch](/news/gemini-3-8-live) for Google's real-time voice models. The voice is cheap now; the distribution is the part worth automating.

Quick takeaways

  • Google launched Gemini 3.8 Flash TTS and Flash-Lite TTS on September 23, 2026, via the Gemini API and Google AI Studio, with Gemini Enterprise to follow.
  • Both cover more than 100 languages, offer 30 studio voices plus a 2,000+ voice library, and support voice design from text, cloning from a 30-second sample, two-speaker dialogue, and line-by-line delivery control.
  • Promotional pricing through end-2026: $0.50/M input tokens plus $9.00/M (Flash) or $6.00/M (Flash-Lite) audio output — roughly $0.81 and $0.54 per hour of speech.
  • Every clip carries an inaudible SynthID watermark; voice cloning requires a recorded consent statement from the person being cloned.

Frequently asked questions

What is Gemini 3.8 text-to-speech?

It's a pair of Google text-to-speech models — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS — launched September 23, 2026, that convert written scripts into expressive spoken audio in more than 100 languages. You can pick from 30 studio voices or a 2,000+ voice library, design a new voice from a text description, clone one from a 30-second sample, stage two-speaker dialogue, and direct pacing and emotion line by line. They're available through the Gemini API and Google AI Studio.

What's the difference between Gemini 3.8 Flash TTS and Flash-Lite TTS?

Flash TTS is the higher-quality tier aimed at creative work — podcasts, audiobooks, game characters — and covers the widest set of languages. Flash-Lite TTS is a cheaper tier built for speech at scale like dubbing, bulk audio, and voice agents, with slightly fewer languages. On promotional pricing through end-2026, audio output is $9.00 per million tokens on Flash and $6.00 on Flash-Lite, or roughly $0.81 versus $0.54 per hour of speech.

Can Gemini 3.8 TTS clone a voice?

Yes. Both models can build a voice profile from a 30-second audio sample, but the person whose voice is being cloned has to record a spoken statement of consent first, and every generated clip carries an inaudible SynthID watermark for AI detection. You can also design an entirely new synthetic voice by describing it in plain text rather than cloning a real person.

How does a creator turn Gemini TTS audio into social content?

The model produces an audio file, not posts. A content engine like Kompozy takes a voiced recording, clips it into captioned vertical shorts, reframes per platform, and generates matching carousels, a blog, a newsletter, and text posts from the same source under one Persona Brief — then schedules and publishes the batch across the eight social platforms plus blog and email. The voice is generated in Gemini; the distribution and multi-format repurposing happen in Kompozy.

Related news

← All AI news · Get started →