// AI NEWS · MODEL RELEASE

Google Launches Gemini 3.5 Transcribe — Speech-to-Text That Strips the "Ums" and "Ahs" Automatically

Google's new speech-to-text model converts raw audio directly into clean, formatted text: it removes filler words, fixes self-corrections, and supports voice editing across 85+ languages. It replaces Chirp 3.

2026-08-26 · by Moe Ameen

What happened

On August 26, 2026, Google introduced Gemini 3.5 Transcribe, its most accurate speech-to-text model to date and the successor to Chirp 3. Instead of writing down every sound literally, it converts raw audio directly into clean, formatted text — automatically removing filler words like "ums" and "ahs," resolving self-corrections, punctuating on its own, and letting you edit by voice.

Google reports an average word error rate of 4.0% for streaming audio and 2.6% for pre-recorded files, with roughly a 70% latency improvement over Chirp 3. The model automatically detects and transcribes across more than 85 languages, handles regional accents, identifies up to three speakers with timestamps, and accepts custom vocabulary so names and jargon come out right.

The rollout spans Google's surfaces. On the consumer side it's in the Gemini app on macOS (English) and it powers Rambler, the dictation feature on Android that debuted on the Pixel 11, with Chrome support coming. Developers get it in public preview through the Gemini API in Google AI Studio and Google Antigravity, and enterprises through the Gemini Enterprise Agent Platform. Google did not publish standalone pricing at launch, and the consumer surfaces are staged by country and language — confirm current availability and rates on Google's own pages before depending on any figure.

Why it matters for creators

  • Clean text from messy speech. A rambled voice memo or a recorded talk comes back already de-ummed and formatted, so the raw material you feed a content pipeline is usable immediately instead of needing a manual cleanup pass.
  • Dictation becomes a real capture method. With filler-word removal and voice editing, thinking out loud into your phone produces a coherent draft rather than a wall of "uh" you then have to rewrite.
  • Speaker labels and timestamps help repurposing. Interview and podcast transcripts that mark who spoke and when are far easier to slice into quotes, clips, and captions.
  • It is a transcript, not a post. The model stops at text — it makes no video, images, feed captions, or scheduled posts. The production and distribution work is entirely untouched.
  • Preview, and vendor-reported numbers. The accuracy stats are Google's own and the consumer surfaces are staged; treat the specifics as early until independently confirmed.

How to act on this with Kompozy

Read this as friction leaving the front of the pipeline. The hardest part of consistent content usually isn't the idea — it's getting the idea out of your head and into a usable form. Gemini 3.5 Transcribe shrinks that step: talk through a topic for two minutes, or transcribe a webinar you already recorded, and you get clean, de-ummed text back. What it can't do is turn that text into anything you can post. That's where [Kompozy](/) picks up.

Drop the cleaned transcript into Kompozy as a source and the engine fans it into roughly 25–35 finished assets across 18 formats — a captioned [Persona Short](/glossary/persona-shorts) with a face-locked avatar delivering your point, brand-exact [carousels](/glossary/hyperframes), quote graphics, photo posts, a blog article, and an email newsletter — each rewritten under a [Persona Brief](/glossary/persona-brief) so the voice stays yours rather than reading like a raw transcript. [Autopilot](/glossary/autopilot) then schedules and publishes the set across the eight social platforms plus blog and email behind a per-post review gate. The workflow a lot of creators land on: dictate or transcribe with Google, generate and ship with Kompozy — a voice memo in the morning becomes a week of posts by lunch.

Quick takeaways

  • Gemini 3.5 Transcribe = Google's new speech-to-text model, launched Aug 26, 2026, replacing Chirp 3.
  • It removes filler words, fixes self-corrections, auto-formats, and supports voice editing; 4.0% streaming / 2.6% non-streaming WER.
  • 85+ languages, up to 3 speakers with timestamps; in the Gemini app, Android Rambler, and the Gemini API (public preview).
  • It produces clean text, not posts — Kompozy turns that transcript into finished, scheduled content across platforms.

Frequently asked questions

What is Gemini 3.5 Transcribe?

It is Google's speech-to-text model announced August 26, 2026, and the successor to Chirp 3. It converts audio into polished, formatted text, automatically removing filler words like "ums" and "ahs," resolving self-corrections, and supporting voice editing across more than 85 languages.

Does Gemini 3.5 Transcribe remove filler words?

Yes. Its "smart transcription" automatically strips filler words such as "ums" and "ahs" and cleans up self-corrections, so the output reads as polished text rather than a literal, word-for-word dictation.

Where can I use Gemini 3.5 Transcribe?

On the consumer side it is in the Gemini app on macOS and powers Rambler dictation on Android and the Pixel 11, with Chrome coming. Developers can access it in public preview through the Gemini API in Google AI Studio and Google Antigravity. Confirm current availability on Google's pages.

Can Gemini 3.5 Transcribe create social media posts?

No. It transcribes speech into clean text and stops there — it makes no video, images, captioned clips, or scheduled posts. To turn a transcript into finished content across platforms you pair it with a generation-and-publishing engine like Kompozy.

Related news

← All AI news · Get started →