Google's new speech-to-text model converts raw audio directly into clean, formatted text: it removes filler words, fixes self-corrections, and supports voice editing across 85+ languages. It replaces Chirp 3.
2026-08-26 · by Moe Ameen
On August 26, 2026, Google introduced Gemini 3.5 Transcribe, its most accurate speech-to-text model to date and the successor to Chirp 3. Instead of writing down every sound literally, it converts raw audio directly into clean, formatted text — automatically removing filler words like "ums" and "ahs," resolving self-corrections, punctuating on its own, and letting you edit by voice.
Google reports an average word error rate of 4.0% for streaming audio and 2.6% for pre-recorded files, with roughly a 70% latency improvement over Chirp 3. The model automatically detects and transcribes across more than 85 languages, handles regional accents, identifies up to three speakers with timestamps, and accepts custom vocabulary so names and jargon come out right.
The rollout spans Google's surfaces. On the consumer side it's in the Gemini app on macOS (English) and it powers Rambler, the dictation feature on Android that debuted on the Pixel 11, with Chrome support coming. Developers get it in public preview through the Gemini API in Google AI Studio and Google Antigravity, and enterprises through the Gemini Enterprise Agent Platform. Google did not publish standalone pricing at launch, and the consumer surfaces are staged by country and language — confirm current availability and rates on Google's own pages before depending on any figure.
Read this as friction leaving the front of the pipeline. The hardest part of consistent content usually isn't the idea — it's getting the idea out of your head and into a usable form. Gemini 3.5 Transcribe shrinks that step: talk through a topic for two minutes, or transcribe a webinar you already recorded, and you get clean, de-ummed text back. What it can't do is turn that text into anything you can post. That's where [Kompozy](/) picks up.
Drop the cleaned transcript into Kompozy as a source and the engine fans it into roughly 25–35 finished assets across 18 formats — a captioned [Persona Short](/glossary/persona-shorts) with a face-locked avatar delivering your point, brand-exact [carousels](/glossary/hyperframes), quote graphics, photo posts, a blog article, and an email newsletter — each rewritten under a [Persona Brief](/glossary/persona-brief) so the voice stays yours rather than reading like a raw transcript. [Autopilot](/glossary/autopilot) then schedules and publishes the set across the eight social platforms plus blog and email behind a per-post review gate. The workflow a lot of creators land on: dictate or transcribe with Google, generate and ship with Kompozy — a voice memo in the morning becomes a week of posts by lunch.
It is Google's speech-to-text model announced August 26, 2026, and the successor to Chirp 3. It converts audio into polished, formatted text, automatically removing filler words like "ums" and "ahs," resolving self-corrections, and supporting voice editing across more than 85 languages.
Yes. Its "smart transcription" automatically strips filler words such as "ums" and "ahs" and cleans up self-corrections, so the output reads as polished text rather than a literal, word-for-word dictation.
On the consumer side it is in the Gemini app on macOS and powers Rambler dictation on Android and the Pixel 11, with Chrome coming. Developers can access it in public preview through the Gemini API in Google AI Studio and Google Antigravity. Confirm current availability on Google's pages.
No. It transcribes speech into clean text and stops there — it makes no video, images, captioned clips, or scheduled posts. To turn a transcript into finished content across platforms you pair it with a generation-and-publishing engine like Kompozy.