The updated model handles batch and streaming audio, adds speaker labels, word-level timestamps, and filler-word removal at no extra cost, and lands its biggest accuracy gain on multilingual short phrases.
2026-09-18 · by Moe Ameen
xAI released Grok Voice Transcribe 2.0 on September 18, 2026, a speech-to-text model available through its Speech-to-Text API. The headline claim is efficiency: xAI describes it as "twice as accurate as Grok Voice Transcribe 1.0, at the same price," and says it improves on the prior version across all of its internal evaluation sets. On Artificial Analysis's leaderboard, xAI says the model ranks first for accuracy among the streaming models tested.
The model handles two modes: batch transcription of recorded files and URLs, and real-time streaming. Priced by audio duration, batch runs $0.10 per hour of audio and streaming $0.20 per hour, with features that other providers often meter separately — speaker diarization and word-level timestamps with confidence scores — included at no additional cost. Other capabilities include support for up to eight independent audio channels, key term biasing of up to 100 domain terms per request, automatic formatting of numbers, dates, currencies, phone numbers, and emails, optional filler-word removal, and smart turn detection aimed at voice agents.
The largest accuracy gain is multilingual. The model transcribes dozens of languages, detects the language automatically, and can follow a mid-recording language switch in a single pass. On xAI's short-phrase set of voice-assistant utterances spanning 19 languages, it reports word error rate dropping from 20.6% to 6.8%, and it says the model leads every model it tested on a telephony set of 8 kHz English customer-support calls.
Existing API integrations receive the upgrade automatically; developers who want to stay on the older behavior can pin grok-voice-transcribe-1.0. As with any fresh release, the accuracy figures are vendor-reported at launch — treat them as a benchmark rather than a guarantee and confirm current pricing and specs against xAI's own documentation.
Read this as a price cut on the front of the content pipeline. The slow, expensive part of repurposing has always been getting raw audio and video into a usable, timestamped form — and Grok Voice Transcribe 2.0 makes that step cheaper and, for multilingual creators, meaningfully more accurate. But a clean transcript with speaker labels and word-level timestamps is not a post. It is raw material. The distance from "I have an accurate transcript" to "on-brand posts scheduled across every platform" is the entire job, and that is where [Kompozy](/) operates.
The concrete loop uses the exact metadata this release gives away for free. Transcribe your recording with word-level timestamps, then bring the source into Kompozy and let the engine do the part a transcription model can't. From one input Kompozy generates roughly 25–35 finished assets across 18 formats — [Clipped Shorts](/glossary/clipped-short) cut straight from the long-form, a captioned [Persona Short](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, brand-exact [Carousel Posts](/glossary/hyperframes), quote graphics pulled from the strongest lines, a blog article, and an email newsletter — each governed by a [Persona Brief](/glossary/persona-brief) so the voice reads as yours, cleared through a per-post review gate, then scheduled and published by [Autopilot](/glossary/autopilot) across the eight social platforms plus blog and email. Cheaper, more accurate transcription is a tailwind for that workflow, not a replacement for it. If you want the step-by-step, see [how to turn a video into clean text](/how-to/turn-a-video-into-clean-text) and then let Kompozy take it the rest of the way.
It is xAI's speech-to-text model, released September 18, 2026, and available through its Speech-to-Text API. It transcribes recorded files, URLs, and real-time streams, and includes speaker diarization and word-level timestamps at no extra cost. xAI describes it as twice as accurate as Grok Voice Transcribe 1.0 at the same price.
xAI prices it by audio duration: $0.10 per hour of audio for batch transcription and $0.20 per hour for streaming. Features like speaker diarization and word-level timestamps are included at no additional cost. Confirm current pricing on xAI's documentation before you build on it.
xAI says it improves on v1.0 across all its internal evaluation sets and ranks first for accuracy among the streaming models tested on Artificial Analysis. Its largest gain is multilingual: on a 19-language short-phrase set, it reports word error rate dropping from 20.6% to 6.8%. Those figures are vendor-reported, so treat them as a benchmark.
No. It produces a transcript with timestamps and speaker labels — it cuts no clips, writes no posts, generates no video or images, and publishes nothing. To turn a recording into finished, scheduled content, you pair it with a content engine like Kompozy.