Google's Gemini 3.8 Flash TTS and Flash-Lite TTS — expressive voice generation with custom voice design, cloning, two-speaker dialogue, and 100+ languages.
Last verified · 2026-09-23 · by Moe Ameen
Gemini 3.8 Text-to-Speech is a pair of Google models — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, launched September 23, 2026 — that turn written scripts into expressive spoken audio. Flash TTS is the higher-quality, creative tier (podcasts, audiobooks, game characters); Flash-Lite TTS is a cheaper tier built for speech at scale like dubbing, bulk narration, and voice agents. Both are available through the Gemini API and Google AI Studio, with Gemini Enterprise access to follow, and both span more than 100 languages and dialects (Flash TTS covers the widest set).
The defining trait is control. Beyond a base of 30 studio voices and a library of more than 2,000 ready-to-use voices — including regional variants like Mexican Spanish, Quebec French, and Scots English — you can design a brand-new voice by describing it in plain text, or clone one from a 30-second sample (the person being cloned must record a spoken consent statement first). You direct delivery line by line for pacing and emotion, script non-verbal cues like laughter and sighs, and stage a native two-speaker conversation from a single script. The models hold a voice steady across long-form audio measured in hours, and Google says Flash TTS topped Hume AI's Voice Design benchmark.
Pricing is metered per token. On promotional pricing through the end of 2026, it runs $0.50 per million input tokens plus $9.00 per million tokens of audio output on Flash TTS and $6.00 on Flash-Lite — roughly $0.81 and $0.54 per hour of speech, since Google counts one second of audio as 25 tokens. Every clip carries an inaudible SynthID watermark for AI detection. It is a voice-generation model, not a content tool: it outputs an audio file and doesn't clip, caption, or publish. Confirm current specs and pricing in Google's docs, since new models change fast.
Here's the pipeline that actually pays off. Say you generate a 20-minute two-speaker podcast episode with Gemini 3.8 Flash TTS, or a chapter of narration for a faceless channel. That's a great audio asset — and it's invisible until it's cut up, captioned, and posted. [Kompozy](/) is the engine that turns one voiced recording into a week of published content. Pair the audio with footage or a persona to make a video, and Kompozy clips the long piece into captioned vertical shorts, auto-reframes each for TikTok, Reels, and Shorts, and pulls the hooks worth posting — then generates the promotional layer the audio can't: Carousel Posts summarizing the episode, Quote Graphics of the best lines, a Blog Article of the transcript, an Email Newsletter, and native Text Posts, all held to one voice by a Persona Brief.
Kompozy doesn't generate the raw voice — its own persona and avatar video runs on HeyGen's native multi-language TTS — so the honest division of labor is: Gemini voices the episode, Kompozy makes it a multi-format, on-brand presence and ships it. Autopilot and a per-post review pipeline schedule and publish the whole batch across the eight social platforms plus blog and email from a single queue, so a single Gemini-voiced recording becomes a full content calendar instead of a file sitting in a folder.
It's a pair of Google models — Gemini 3.8 Flash TTS and Flash-Lite TTS, launched September 23, 2026 — that convert written scripts into expressive spoken audio in more than 100 languages, with custom voice design, cloning from a 30-second sample, two-speaker dialogue, and line-by-line delivery control. They're available via the Gemini API and Google AI Studio.
Flash TTS is the higher-quality tier for creative work like podcasts, audiobooks, and game characters, with the widest language coverage. Flash-Lite TTS is a cheaper tier for speech at scale — dubbing, bulk audio, voice agents — with slightly fewer languages. On promotional pricing, audio output is $9.00 per million tokens on Flash and $6.00 on Flash-Lite, or roughly $0.81 vs $0.54 per hour.
Yes. It supports native two-speaker staging from a single script, so you can generate a back-and-forth conversation between two distinct voices without recording or booking a second person. You can also direct pacing and emotion line by line and add non-verbal cues like laughter.
Kompozy's own persona and avatar video uses HeyGen's native multi-language text-to-speech, not the Gemini TTS API. The bridge is that you generate a voiced recording in Gemini, then bring it into Kompozy to clip, caption, reframe, and repurpose into carousels, blogs, and newsletters, and publish it across platforms.
The model outputs an audio file only. Kompozy takes a voiced recording, clips it into captioned vertical shorts, reframes per platform, and generates matching carousels, quote graphics, a blog, a newsletter, and text posts from the same source — then schedules and publishes them across nine platforms.