// AI NEWS · MODEL RELEASE

xAI Ships Grok Voice Transcribe 2.0, a Speech-to-Text Model It Says Is Twice as Accurate as v1.0 at the Same Price

The updated model handles batch and streaming audio, adds speaker labels, word-level timestamps, and filler-word removal at no extra cost, and lands its biggest accuracy gain on multilingual short phrases.

2026-09-18 · by Moe Ameen

What happened

xAI released Grok Voice Transcribe 2.0 on September 18, 2026, a speech-to-text model available through its Speech-to-Text API. The headline claim is efficiency: xAI describes it as "twice as accurate as Grok Voice Transcribe 1.0, at the same price," and says it improves on the prior version across all of its internal evaluation sets. On Artificial Analysis's leaderboard, xAI says the model ranks first for accuracy among the streaming models tested.

The model handles two modes: batch transcription of recorded files and URLs, and real-time streaming. Priced by audio duration, batch runs $0.10 per hour of audio and streaming $0.20 per hour, with features that other providers often meter separately — speaker diarization and word-level timestamps with confidence scores — included at no additional cost. Other capabilities include support for up to eight independent audio channels, key term biasing of up to 100 domain terms per request, automatic formatting of numbers, dates, currencies, phone numbers, and emails, optional filler-word removal, and smart turn detection aimed at voice agents.

The largest accuracy gain is multilingual. The model transcribes dozens of languages, detects the language automatically, and can follow a mid-recording language switch in a single pass. On xAI's short-phrase set of voice-assistant utterances spanning 19 languages, it reports word error rate dropping from 20.6% to 6.8%, and it says the model leads every model it tested on a telephony set of 8 kHz English customer-support calls.

Existing API integrations receive the upgrade automatically; developers who want to stay on the older behavior can pin grok-voice-transcribe-1.0. As with any fresh release, the accuracy figures are vendor-reported at launch — treat them as a benchmark rather than a guarantee and confirm current pricing and specs against xAI's own documentation.

Why it matters for creators

  • Cheap, accurate transcription lowers the cost of the first step in repurposing. At $0.10 per hour of audio for batch, turning an hour-long podcast or interview into clean text is close to free — the bottleneck moves downstream to production.
  • The multilingual jump is the most useful part for creators. Auto-detection plus mid-recording language switching means creators filming in more than one language get a usable transcript in one pass, no manual language tagging.
  • Word-level timestamps and speaker labels come free. That metadata is exactly what a clipping workflow needs to find the strong moments and attribute quotes — it is more valuable than the raw transcript for anyone repurposing long video.
  • It is a transcription model, not a content tool. It returns text with metadata; it cuts no clips, writes no posts, makes no video or images, and publishes nothing. The transcript is an ingredient, not the finished post.
  • Much of the feature set — streaming, turn detection, telephony tuning, eight audio channels — is aimed at voice agents and call centers, not creators. A creator uses the batch mode and the timestamps; the rest is for developers.

How to act on this with Kompozy

Read this as a price cut on the front of the content pipeline. The slow, expensive part of repurposing has always been getting raw audio and video into a usable, timestamped form — and Grok Voice Transcribe 2.0 makes that step cheaper and, for multilingual creators, meaningfully more accurate. But a clean transcript with speaker labels and word-level timestamps is not a post. It is raw material. The distance from "I have an accurate transcript" to "on-brand posts scheduled across every platform" is the entire job, and that is where [Kompozy](/) operates.

The concrete loop uses the exact metadata this release gives away for free. Transcribe your recording with word-level timestamps, then bring the source into Kompozy and let the engine do the part a transcription model can't. From one input Kompozy generates roughly 25–35 finished assets across 18 formats — [Clipped Shorts](/glossary/clipped-short) cut straight from the long-form, a captioned [Persona Short](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, brand-exact [Carousel Posts](/glossary/hyperframes), quote graphics pulled from the strongest lines, a blog article, and an email newsletter — each governed by a [Persona Brief](/glossary/persona-brief) so the voice reads as yours, cleared through a per-post review gate, then scheduled and published by [Autopilot](/glossary/autopilot) across the eight social platforms plus blog and email. Cheaper, more accurate transcription is a tailwind for that workflow, not a replacement for it. If you want the step-by-step, see [how to turn a video into clean text](/how-to/turn-a-video-into-clean-text) and then let Kompozy take it the rest of the way.

Quick takeaways

  • xAI released Grok Voice Transcribe 2.0 on September 18, 2026, calling it twice as accurate as v1.0 at the same price.
  • Pricing is $0.10 per hour of audio for batch and $0.20 per hour for streaming; speaker diarization and word-level timestamps are included free.
  • It transcribes dozens of languages with auto-detection and mid-recording switching; the biggest gain is multilingual short phrases, with reported word error rate dropping from 20.6% to 6.8%.
  • Existing API integrations upgrade automatically; developers can pin grok-voice-transcribe-1.0 to stay on the old model.
  • It produces text plus metadata, not finished content — Kompozy turns that transcript into clips, posts, video, and a published multi-platform package.

Frequently asked questions

What is Grok Voice Transcribe 2.0?

It is xAI's speech-to-text model, released September 18, 2026, and available through its Speech-to-Text API. It transcribes recorded files, URLs, and real-time streams, and includes speaker diarization and word-level timestamps at no extra cost. xAI describes it as twice as accurate as Grok Voice Transcribe 1.0 at the same price.

How much does Grok Voice Transcribe 2.0 cost?

xAI prices it by audio duration: $0.10 per hour of audio for batch transcription and $0.20 per hour for streaming. Features like speaker diarization and word-level timestamps are included at no additional cost. Confirm current pricing on xAI's documentation before you build on it.

How accurate is Grok Voice Transcribe 2.0?

xAI says it improves on v1.0 across all its internal evaluation sets and ranks first for accuracy among the streaming models tested on Artificial Analysis. Its largest gain is multilingual: on a 19-language short-phrase set, it reports word error rate dropping from 20.6% to 6.8%. Those figures are vendor-reported, so treat them as a benchmark.

Can Grok Voice Transcribe 2.0 create social media content?

No. It produces a transcript with timestamps and speaker labels — it cuts no clips, writes no posts, generates no video or images, and publishes nothing. To turn a recording into finished, scheduled content, you pair it with a content engine like Kompozy.

Related news

← All AI news · Get started →