Edit X's auto-generated video captions by hand, fix what transcription breaks, and use Grok to correct and translate them — the 2026 editable-captions workflow.
Last verified · 2026-10-06 · by Moe Ameen
X made its video captions editable in early October 2026. Before the update, the captions X auto-generated from your clip's audio were effectively locked — a wrong brand name or a mangled number stayed wrong. Now, inside the video composer, you can edit the generated transcript by hand, ask Grok to correct it, or tell Grok exactly how to change it, including translating it into another language. The announcement came from Allegra Jacchia, who leads X's Creators Product team.
This is the caption-editing task specifically — getting the words right and, if you need it, into a second language — not a tour of the whole composer (for that, see [how to use X's updated video composer](/how-to/use-x-updated-video-composer)). The steps below are the lifecycle of a caption on X: generate it, read it against the audio, fix the predictable errors, hand the hard cases to Grok, translate and verify, then check it reads on a muted timeline before you post. The whole point of editable captions is that there is no longer an excuse to ship the raw transcription — so the real work is the review, and that is what this covers.
Read back through these steps and notice what the task actually is: proofreading a machine's guess at what your audio said. You generate a transcription, then hunt for the names, numbers, and homophones it got wrong, and verify a translation you may not be able to read. X's editable captions make that hunt faster, but the hunt is still the job — because the caption is reverse-engineered from audio, so the errors are baked in before you ever open the composer. [Kompozy](/) attacks the problem from the other end: it removes the thing you are proofreading.
In Kompozy the caption is not a transcription of audio — it is the authored script the video was built from, so the words on screen are the words you wrote, not a model's best guess at them. That flips the error class this tutorial is all about. Your brand and product names are spelled the way the [Persona Brief](/glossary/persona-brief) says to spell them, so "my company is spelled Kompozy, fix every instance" never has to be typed because it was never mis-transcribed. Numbers come from the source copy, not from a model mishearing a price. The captions are then burned into the frame, word-synced, from a reusable style preset in HyperFrames — so placement and the sound-off read are set once instead of re-checked clip by clip. For a genuine second-language version, a HeyGen persona can speak the target language on camera rather than only subtitling it, which is the one thing Grok's translation explicitly does not do.
Honest boundary: to fix the captions on one clip you already filmed for X, use X's composer — it is right there and it is good at it. Kompozy is for when captioned, correctly-spelled, on-brand video has to ship across the eight social platforms plus blog and email on a repeating cadence, and the hand-proofing you just did does not scale to that. Starter ($199/mo for 5,500 credits) fits a solo creator; Pro ($499/mo for 18,000 credits) suits a brand or agency running many accounts; Enterprise is custom.
Yes. As of an early-October 2026 update to X's native video composer, you can edit the auto-generated caption transcript by hand before you post, rather than accepting the raw transcription. You can also ask Grok to correct the captions or tell it exactly how to change them. The fix is part of the upload, so do it before posting.
In the updated composer, ask Grok to translate the caption track into your target language. This is X's built-in path to subtitling a clip for a non-English audience without a separate tool. It translates the on-screen caption text only — it does not dub or re-voice the spoken audio — so review the result for names and idioms before posting.
Just the captions. Grok translates the caption text that appears on screen; the video's spoken audio stays in its original language. If you need the voice itself to speak another language, that requires dubbing or re-voicing, which X's caption tools do not do — they subtitle, not dub.
Proper nouns and numbers. Auto-transcription handles plain speech well but reliably stumbles on brand and product names, people's names, prices and figures, acronyms, and homophones (their/there, to/two). Those are the lines to read against the audio and fix first, because they are the errors that change meaning or make a caption look careless.
No. The captions, trims, and styling you create in X's composer are baked into that X upload only. Download the file and post it to Instagram, TikTok, or YouTube and none of the X-native editing travels — you caption and crop again for each platform. That per-platform repetition is the real cost of relying on each app's native editor.