Gemini 3.8 TTS review (2026): honest scoring on voice quality, cloning, two-speaker dialogue, 100+ languages, token pricing, and where the workflow stops.
Gemini 3.8 Flash TTS and Flash-Lite TTS are among the best-value expressive voice models available: design a voice from a text description, clone one from 30 seconds, stage two-speaker dialogue from a single script, and generate hours of stable audio in 100+ languages for well under a dollar an hour. As a voice engine it earns a high score. The honest ceiling is scope — it outputs an audio file and nothing else. There is no clipping, captioning, other formats, or publishing, so a finished voice track is the start of a content workflow, not the end.
Google's Gemini 3.8 text-to-speech models — Flash TTS and Flash-Lite TTS, launched September 23, 2026 — are the company's most expressive voice models to date, and this review scores them for the creator who searches "Gemini 3.8 TTS review" wondering whether to build narration, podcasts, or dubbed audio on them. The lens here is "does this help me produce and ship content," not "is this the best raw API for a voice agent," though it happens to be strong at both.
The capability set is deep. You get a base of 30 studio voices plus a library of more than 2,000 ready-to-use voices with regional variants, and two ways to go beyond them: design a brand-new voice by describing it in plain text, or clone an existing one from a 30-second sample (with a recorded consent statement required). You can direct delivery line by line for pacing and emotion, script non-verbal cues like laughter and sighs, and stage a two-speaker conversation from a single script — and the models hold a voice steady across long-form audio measured in hours. Google says Flash TTS topped Hume AI's Voice Design benchmark, and both models span more than 100 languages.
I score it on the dimensions that fit a voice model: voice quality, voice design, cloning fidelity, language coverage, multi-speaker handling, delivery control, long-form stability, and value — where it does well — plus content workflow, where a raw TTS model necessarily scores low because it produces an audio file and stops. The two-tier structure matters to the score: Flash TTS is the creative, higher-quality tier; Flash-Lite TTS is a cheaper tier for speech at scale.
Everything below reflects the models' launch-window state as of 2026-09-23, verified against Google's announcement and reporting on the release. New models change fast — confirm current pricing, language counts, and features in Google's docs before committing.
Gemini 3.8 text-to-speech is a pair of Google models that convert written scripts into spoken audio. Gemini 3.8 Flash TTS is the higher-quality, more expressive tier aimed at creative work such as podcasts, audiobooks, and game characters; Gemini 3.8 Flash-Lite TTS is a lower-cost tier built for speech at scale like dubbing, bulk audio narration, and voice agents. Both are available through the Gemini API and Google AI Studio, with Gemini Enterprise access to follow, and both cover more than 100 languages and dialects (Flash TTS the widest set). Beyond preset voices, the models let you design a voice from a text description, clone one from a 30-second sample, control delivery line by line, and stage native two-speaker dialogue from a single script — with non-verbal sounds like laughter and sighs. Output is standard audio (WAV, 24 kHz), and every clip carries an inaudible SynthID watermark for AI detection. What it is not is a content tool: it renders audio, and does not clip, caption, reframe, generate other formats, or schedule and publish. Distribution is another product's job.
Gemini 3.8 TTS fits anyone whose deliverable is high-quality spoken audio and who is comfortable working through an API or AI Studio: developers building voice agents and apps, podcasters and audiobook producers who want a directable narrator, game studios voicing characters, and localization teams dubbing at scale in many languages. If you need an expressive, controllable, cheap voice track — especially multilingual or multi-speaker — it is an excellent, well-priced pick. It is a weaker fit for a creator whose actual goal is a consistent, multi-format social presence: the models give you the voice, but nothing to clip it into shorts, caption for feeds, turn into carousels or a blog, or schedule and publish. Non-developers should also note the access path leans technical — this is a model and an API, not a polished consumer editor.
| Dimension | Score | Why |
|---|---|---|
| Voice quality & expressiveness | 4.7 / 5 | Natural, expressive delivery with directable emotion and non-verbal cues; Google reports a top finish on Hume AI's Voice Design benchmark. |
| Voice design (create from text) | 4.5 / 5 | Describe a voice in plain language and get a usable new voice — a genuinely flexible way to build a custom narrator without a real sample. |
| Voice cloning fidelity | 4.3 / 5 | Builds a voice profile from a 30-second sample; the required recorded consent statement is a responsible guardrail, though it adds a step. |
| Language & accent coverage | 4.6 / 5 | 100+ languages and dialects with regional variants like Mexican Spanish and Quebec French — deep reach for localization. |
| Multi-speaker & dialogue | 4.4 / 5 | Native two-speaker staging from a single script makes solo-produced podcasts and conversational explainers practical. |
| Delivery control | 4.5 / 5 | Line-by-line direction of pacing and emotion, plus scripted vocal bursts and backchanneling — real control, not just a monotone read. |
| Long-form stability | 4.3 / 5 | Holds a voice consistent across hours of audio, which matters for audiobooks and long podcasts. |
| Pricing & value | 4.6 / 5 | Roughly $0.81/hr (Flash) and $0.54/hr (Flash-Lite) on promotional pricing — among the best value in expressive TTS. |
| Access & ease of use | 3.8 / 5 | Gemini API and AI Studio are developer-oriented; there is no polished consumer editor, so non-technical creators face a learning curve. |
| Content workflow (clip/caption/publish) | 1.5 / 5 | Produces an audio file and stops — no clipping, captioning, other formats, scheduling, or publishing. |
Gemini 3.8 TTS is priced as a metered API, not a flat subscription, and on that basis it is aggressively cheap. Through the end of 2026, promotional pricing runs $0.50 per million input (text) tokens plus $9.00 per million tokens of audio output on Flash TTS and $6.00 on Flash-Lite TTS. Because Google counts one second of generated audio as 25 tokens, an hour of speech works out to roughly $0.81 on Flash and $0.54 on Flash-Lite. Note the word promotional: the standard rate is set to be about double once the introductory window closes, so plan a real budget on the higher number.
For the job it is built for, this is strong value. An hour of directable, multilingual narration for well under a dollar undercuts hiring voice talent and beats or matches most expressive TTS competitors on a per-minute basis, and the two-tier split lets you spend Flash quality only where it matters and run bulk dubbing or voice-agent traffic on Flash-Lite. Token metering also scales cleanly — you pay for what you generate, with no seat minimums.
The honest caveat is what the price does not include. It buys audio, and only audio. There is no clipping, captioning, reframing, format diversification, or publishing in the number, so if your goal is a content presence rather than a voice track, the model is one input cost among several, not the whole bill. Read it per job: excellent value for generating speech, no coverage for operationalizing it into posts.
| Use case | Fit | Why |
|---|---|---|
| Narrating faceless videos and shorts | Strong | Cheap, expressive, directable narration in many languages is exactly what this is built to produce. |
| Producing AI podcasts with two speakers | Strong | Native two-speaker staging from a single script makes a solo-produced dialogue podcast practical. |
| Audiobooks and long-form narration | Strong | Long-form stability holds one voice consistent across hours, and per-hour cost is very low. |
| Multilingual dubbing and localization at scale | Strong | 100+ languages plus cheap Flash-Lite output make bulk dubbing economical. |
| Voicing game characters or apps | OK | Voice design and cloning are strong, but you build the integration yourself through the API. |
| Turning a voiced recording into per-platform short-form | Weak | The model outputs audio only — no clipping into shorts or per-platform reframing. |
| Building a multi-format content week | Weak | It generates the voice, not the carousels, images, blog, or newsletter around it. |
| Scheduling and publishing to social channels | Weak | There is no scheduler and no publishing — the workflow ends at an audio file. |
Gemini 3.8 TTS and Kompozy aren't rivals — they sit on opposite ends of the same pipeline, and it's more honest to say that than to stage a head-to-head. On the voice itself — expressiveness, cloning, language reach, price — Google is excellent, and Kompozy doesn't try to compete: Kompozy's own video uses HeyGen's native multi-language TTS for talking-head personas, not a general-purpose voice API you script arbitrary audio with. If your task is "generate a great voice track," build on Gemini and I'll say so plainly.
Where Kompozy takes over is everything after the audio renders. A voiced podcast, narration, or dubbed track is an asset, not a content presence. Bring that recording into Kompozy and it clips the long audio-backed video into captioned vertical shorts, reframes each for TikTok, Reels, and Shorts, and spins the same source into Carousel Posts, Photo Posts, a Blog Article, an Email Newsletter, and native Text Posts — all governed by a Persona Brief so voice stays consistent across formats. Autopilot and a per-post review pipeline then schedule and publish the set across the eight social platforms plus blog and email. The clean division: Gemini to voice it, Kompozy to turn each voiced piece into a published, multi-format stream.
For generating expressive, controllable speech, yes — it is one of the best-value voice models available, with voice design from text, cloning from a 30-second sample, two-speaker dialogue, 100+ languages, and roughly $0.54–$0.81 per hour of audio on promotional pricing. The caveat is scope: it produces an audio file and nothing else, so if your goal is published content rather than a voice track, budget for a separate stack to clip, caption, and distribute it.
It is metered per token, not sold as a flat plan. Through end-2026, promotional pricing is $0.50 per million input tokens plus $9.00 per million tokens of audio output on Flash TTS and $6.00 on Flash-Lite. At 25 tokens per second of audio, that is roughly $0.81/hr and $0.54/hr respectively. The standard rate is set to be about double after the promotional window, so plan on the higher figure.
Flash TTS is the higher-quality, more expressive tier for creative work — podcasts, audiobooks, game characters — and covers the widest set of languages. Flash-Lite TTS is a cheaper tier for speech at scale such as dubbing, bulk audio, and voice agents, with slightly fewer languages and a lower per-hour cost.
Yes. It can build a voice profile from a 30-second sample, but the person being cloned must record a spoken statement of consent first, and every generated clip carries an inaudible SynthID watermark. You can also design a new synthetic voice from a text description instead of cloning a real one.
They're close peers on expressive TTS and cloning. Gemini competes hard on price and language breadth and is tightly integrated with Google's stack; ElevenLabs has a more consumer-friendly studio and a long track record with creators. For most people the deciding factor is the surrounding workflow — neither clips, captions, or publishes — so the bigger question is what turns the audio into posts.
No. It renders an audio file and stops; there is no clipping, captioning, reframing, other formats, scheduling, or publishing. A content engine like Kompozy takes the voiced recording and turns it into captioned shorts, carousels, a blog, a newsletter, and text posts, then schedules and publishes them across the eight social platforms plus blog and email.
It is usable through Google AI Studio without heavy coding, but access still leans technical — it is a model and an API, not a polished consumer video or audio editor. Non-technical creators can generate audio, but building it into finished, published content is where a purpose-built content tool helps.
See Gemini 3.8 Text-to-Speech vs Kompozy comparison → · Get Started →