Google's 2026 speech-to-text model that turns raw audio into clean, formatted text — automatically removing filler words like "ums" and "ahs" and fixing self-corrections.
Last verified · 2026-08-26 · by Moe Ameen
Gemini 3.5 Transcribe is Google's speech-to-text model, announced August 26, 2026, as the successor to Chirp 3. Rather than transcribing every sound literally, it converts raw audio directly into clean, formatted text: it removes filler words like "ums" and "ahs," resolves self-corrections, punctuates and formats on its own, and supports editing by voice.
Google reports an average word error rate of 4.0% on streaming audio and 2.6% on pre-recorded files, with roughly a 70% latency improvement over Chirp 3. It automatically detects and transcribes more than 85 languages, handles regional accents, labels up to three speakers with timestamps, and accepts custom vocabulary so names and jargon come out right.
You reach it several ways. Consumers get it in the Gemini app on macOS and through Rambler, the dictation feature on Android that debuted on the Pixel 11, with Chrome support coming. Developers can call it in public preview through the Gemini API in Google AI Studio and Google Antigravity, and enterprises through the Gemini Enterprise Agent Platform. Standalone pricing wasn't published at launch, so confirm current rates and availability on Google's own pages.
For a creator, its value sits at the top of the funnel: it's the fastest way to get a clean, quotable transcript out of a voice memo, an interview, a webinar, or a long video. It writes no posts and renders no media — it's an input tool, and a very good one.
Every content engine needs a source, and a transcript is one of the best ones. Gemini 3.5 Transcribe gives you two clean front doors into a pipeline: dictate a fresh idea and get formatted, de-ummed text back, or run an existing recording — a podcast episode, a client webinar, a talk you already gave — through it and get a quotable transcript with speaker labels. Neither of those is a post yet. [Kompozy](/) is what converts either one into a week of finished, published content.
Feed a dictated transcript into Kompozy and it drafts under your [Persona Brief](/glossary/persona-brief) so a spoken ramble becomes on-brand copy, then fans it into a [Persona Short](/glossary/persona-shorts), [carousels](/glossary/hyperframes), quote graphics, a blog article, and a newsletter. Feed it a long recording instead and Kompozy leans on its own clipping — it cuts [Clipped Shorts](/glossary/clipped-short) from the video and burns in captions, using the transcript as the map of where the good moments are. Either way, [Autopilot](/glossary/autopilot) reframes each output to 9:16, 1:1, and 16:9 and publishes across the eight social platforms plus blog and email behind a per-post review gate. Google cleans the words; Kompozy makes the media and ships it.
Google's speech-to-text model, announced August 26, 2026, and the successor to Chirp 3. It converts audio into clean, formatted text, automatically removing filler words like "ums" and "ahs," resolving self-corrections, and supporting voice editing across 85+ languages.
Google reports an average word error rate of 4.0% on streaming audio and 2.6% on pre-recorded files, with roughly a 70% latency improvement over Chirp 3. Real-world accuracy varies with audio quality, so treat the vendor figures as a benchmark rather than a guarantee.
Consumers use it in the Gemini app on macOS and through Rambler dictation on Android and the Pixel 11, with Chrome coming. Developers call it in public preview via the Gemini API in Google AI Studio and Google Antigravity. Confirm current availability on Google's pages.
No. It produces clean text — a transcript or dictation — and nothing more. It renders no video or images and publishes nothing. To turn a transcript into finished posts across platforms, pair it with a content engine like Kompozy.
It's a strong first step. A clean, speaker-labeled transcript with timestamps is exactly what makes an interview or webinar easy to slice into clips and quotes. The slicing, captioning, and publishing still happen in a separate tool such as Kompozy.