xAI's updated speech-to-text model, released September 18, 2026 — batch and real-time streaming transcription that xAI says is twice as accurate as v1.0 at the same price, with speaker labels and word-level timestamps included at no extra cost.
Last verified · 2026-09-18 · by Moe Ameen
Grok Voice Transcribe 2.0 is xAI's speech-to-text model, released on September 18, 2026, and reachable through its Speech-to-Text API. It converts audio to text in two modes: batch transcription of recorded files and URLs, and real-time streaming. The headline claim is efficiency — xAI describes it as "twice as accurate as Grok Voice Transcribe 1.0, at the same price" and says it improves on the prior version across all of its internal evaluation sets, ranking first for accuracy among the streaming models tested on Artificial Analysis.
The biggest gain is multilingual. It transcribes dozens of languages, detects the language automatically, and can follow a mid-recording language switch in a single pass. On xAI's short-phrase set of voice-assistant utterances spanning 19 languages, it reports word error rate dropping from 20.6% to 6.8%, and it says the model leads every model it tested on a telephony set of 8 kHz English customer-support calls.
Beyond the raw transcript, it returns word-level timestamps with confidence scores and speaker diarization at no additional cost, plus support for up to eight independent audio channels, key term biasing of up to 100 domain terms per request, automatic formatting for numbers, dates, currencies, phone numbers, and emails, optional filler-word removal, and smart turn detection aimed at voice agents. Pricing is by audio duration: $0.10 per hour of audio for batch and $0.20 for streaming. Existing API integrations upgrade automatically, and developers can pin grok-voice-transcribe-1.0 to stay on the old model.
The clean framing for a creator: this is a transcription engine, not a content engine. It listens and writes structured text with metadata; it renders no video, cuts no clips, and publishes nothing. As a fresh release, its figures are vendor-reported — confirm current pricing and specs on xAI's documentation.
The single most useful thing Grok Voice Transcribe 2.0 hands a creator isn't the transcript — it's the metadata attached to it. Word-level timestamps with confidence scores tell you exactly where every strong line lands, and speaker diarization tells you who said it. That is precisely the map a repurposing engine needs to cut good clips and attribute quotes correctly, and it is the map most creators never build because scrubbing a recording by hand is slow. But a timestamped, speaker-labeled transcript is still not content. Nothing is cut, designed, branded, or posted. Turning that structured text into finished, on-brand posts across platforms is the job [Kompozy](/) does.
Here is the concrete loop. Transcribe your recording with Grok Voice Transcribe 2.0, then bring the source and its transcript into Kompozy and pick your formats. From that one input Kompozy generates the assets a transcription model can't: [Clipped Shorts](/glossary/clipped-short) cut straight from the long-form at the strong moments, a captioned [Persona Short](/glossary/persona-shorts) fronted by a face-locked HeyGen avatar, brand-exact [Carousel Posts](/glossary/hyperframes), quote graphics pulled from the highest-confidence lines, photo posts, a blog article, and an email newsletter — each rewritten under a [Persona Brief](/glossary/persona-brief) so the voice reads as yours. For bilingual creators the multilingual detection matters twice over: transcribe a two-language stream in one pass, then let Kompozy fan it into localized posts. [Autopilot](/glossary/autopilot) then schedules and publishes the batch across the eight social platforms plus blog and email, every asset clearing a per-post review gate first. On the Founding tier you can bring your own keys so the ingestion layer runs at cost inside Kompozy.
Grok Voice Transcribe 2.0 is xAI's speech-to-text model, released September 18, 2026, and available through its Speech-to-Text API. It transcribes recorded files, URLs, and real-time streams across dozens of languages, and returns text with word-level timestamps and speaker labels. xAI describes it as twice as accurate as Grok Voice Transcribe 1.0 at the same price.
xAI prices it by audio duration: $0.10 per hour of audio for batch transcription and $0.20 per hour for streaming. Features like speaker diarization and word-level timestamps are included at no additional cost. Confirm current pricing on xAI's documentation before building on it.
No. It produces a transcript with timestamps and speaker labels — it cuts no clips, writes no posts, and generates no video or images, and it publishes nothing. To turn what it transcribes into finished, published content, you pair it with a generation-and-publishing engine like Kompozy.
Transcribe the recording, then bring the source into Kompozy. Kompozy uses the timestamps to cut Clipped Shorts, generates a persona/avatar short, carousels, quote graphics, a blog, and a newsletter from the same input under one Persona Brief, and schedules and publishes the set across the eight social platforms plus blog and email.