// AI TOOLS · GEMINI 3.8 TEXT-TO-SPEECH

Gemini 3.8 Text-to-Speech

Google's Gemini 3.8 Flash TTS and Flash-Lite TTS — expressive voice generation with custom voice design, cloning, two-speaker dialogue, and 100+ languages.

Last verified · 2026-09-23 · by Moe Ameen

What Gemini 3.8 Text-to-Speech is

Gemini 3.8 Text-to-Speech is a pair of Google models — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, launched September 23, 2026 — that turn written scripts into expressive spoken audio. Flash TTS is the higher-quality, creative tier (podcasts, audiobooks, game characters); Flash-Lite TTS is a cheaper tier built for speech at scale like dubbing, bulk narration, and voice agents. Both are available through the Gemini API and Google AI Studio, with Gemini Enterprise access to follow, and both span more than 100 languages and dialects (Flash TTS covers the widest set).

The defining trait is control. Beyond a base of 30 studio voices and a library of more than 2,000 ready-to-use voices — including regional variants like Mexican Spanish, Quebec French, and Scots English — you can design a brand-new voice by describing it in plain text, or clone one from a 30-second sample (the person being cloned must record a spoken consent statement first). You direct delivery line by line for pacing and emotion, script non-verbal cues like laughter and sighs, and stage a native two-speaker conversation from a single script. The models hold a voice steady across long-form audio measured in hours, and Google says Flash TTS topped Hume AI's Voice Design benchmark.

Pricing is metered per token. On promotional pricing through the end of 2026, it runs $0.50 per million input tokens plus $9.00 per million tokens of audio output on Flash TTS and $6.00 on Flash-Lite — roughly $0.81 and $0.54 per hour of speech, since Google counts one second of audio as 25 tokens. Every clip carries an inaudible SynthID watermark for AI detection. It is a voice-generation model, not a content tool: it outputs an audio file and doesn't clip, caption, or publish. Confirm current specs and pricing in Google's docs, since new models change fast.

What you can make with it

  • Expressive narration for faceless videos, explainers, and shorts in 100+ languages
  • Two-speaker AI podcasts and conversational episodes generated from a single script
  • Audiobook and long-form narration held to one consistent voice across hours
  • A custom brand voice designed from a text description, or a cloned voice from a 30-second sample
  • Multilingual dubbing and bulk voice-over for localization at low per-hour cost
  • Character and voice-agent audio via the Gemini API

How Kompozy turns Gemini 3.8 Text-to-Speech output into content

Here's the pipeline that actually pays off. Say you generate a 20-minute two-speaker podcast episode with Gemini 3.8 Flash TTS, or a chapter of narration for a faceless channel. That's a great audio asset — and it's invisible until it's cut up, captioned, and posted. [Kompozy](/) is the engine that turns one voiced recording into a week of published content. Pair the audio with footage or a persona to make a video, and Kompozy clips the long piece into captioned vertical shorts, auto-reframes each for TikTok, Reels, and Shorts, and pulls the hooks worth posting — then generates the promotional layer the audio can't: Carousel Posts summarizing the episode, Quote Graphics of the best lines, a Blog Article of the transcript, an Email Newsletter, and native Text Posts, all held to one voice by a Persona Brief.

Kompozy doesn't generate the raw voice — its own persona and avatar video runs on HeyGen's native multi-language TTS — so the honest division of labor is: Gemini voices the episode, Kompozy makes it a multi-format, on-brand presence and ships it. Autopilot and a per-post review pipeline schedule and publish the whole batch across the eight social platforms plus blog and email from a single queue, so a single Gemini-voiced recording becomes a full content calendar instead of a file sitting in a folder.

  1. Write your script and generate the audio in Gemini 3.8 TTS — a two-speaker podcast, an audiobook chapter, or a faceless-video narration, in your target language.
  2. Turn the audio into a video (pair it with footage, a waveform, or a persona) so it becomes a source Kompozy can mine.
  3. Add that recording to Kompozy as a source and let it clip the long piece into captioned vertical shorts, reframed per platform.
  4. Generate the promo layer from the same source — Carousels, Quote Graphics, a Blog Article, a Newsletter, and Text Posts — under one Persona Brief.
  5. Schedule and publish the batch across nine platforms via Autopilot, approving each post in the review pipeline first.

Frequently asked questions

What is Gemini 3.8 text-to-speech?

It's a pair of Google models — Gemini 3.8 Flash TTS and Flash-Lite TTS, launched September 23, 2026 — that convert written scripts into expressive spoken audio in more than 100 languages, with custom voice design, cloning from a 30-second sample, two-speaker dialogue, and line-by-line delivery control. They're available via the Gemini API and Google AI Studio.

What's the difference between Flash TTS and Flash-Lite TTS?

Flash TTS is the higher-quality tier for creative work like podcasts, audiobooks, and game characters, with the widest language coverage. Flash-Lite TTS is a cheaper tier for speech at scale — dubbing, bulk audio, voice agents — with slightly fewer languages. On promotional pricing, audio output is $9.00 per million tokens on Flash and $6.00 on Flash-Lite, or roughly $0.81 vs $0.54 per hour.

Can Gemini 3.8 TTS make a two-speaker podcast?

Yes. It supports native two-speaker staging from a single script, so you can generate a back-and-forth conversation between two distinct voices without recording or booking a second person. You can also direct pacing and emotion line by line and add non-verbal cues like laughter.

Does Kompozy use Gemini 3.8 TTS?

Kompozy's own persona and avatar video uses HeyGen's native multi-language text-to-speech, not the Gemini TTS API. The bridge is that you generate a voiced recording in Gemini, then bring it into Kompozy to clip, caption, reframe, and repurpose into carousels, blogs, and newsletters, and publish it across platforms.

How do I turn Gemini TTS audio into social posts?

The model outputs an audio file only. Kompozy takes a voiced recording, clips it into captioned vertical shorts, reframes per platform, and generates matching carousels, quote graphics, a blog, a newsletter, and text posts from the same source — then schedules and publishes them across nine platforms.

Related tools

  • Gemini 3.8 LiveGoogle's real-time voice models — a cheap, fast conversational tier plus a reasoning-heavy Extended Thinking variant that thinks and speaks at once, with visual grounding and 97-language switching.
  • SpeechifyA text-to-speech platform built around low-latency streaming voice — its Simba models turn any script into natural narration for reading, voiceover, and developer apps.
  • Kokoro TTSAn open-weight, 82-million-parameter text-to-speech model that runs high-quality narration locally on a CPU — free, offline, and Apache-2.0 licensed for commercial use.
  • Nari Qwen3-TTS & Qwen3-ASRNari Labs' low-latency, low-cost voice models — Qwen3-TTS 1.7B for text-to-speech and Qwen3-ASR 1.7B for transcription — that topped Coval's public voice-AI benchmark in September 2026.

← All AI tools · Get started →