How to use AI-powered auto-caption correction (2026)
Fix auto-captions with AI: keyterm prompting for brand names, an LLM post-editing pass over the ASR transcript, and a glossary that catches the repeat errors.
AI-powered auto-caption correction is the layer that fixes automatic speech recognition (ASR) output automatically, instead of you hand-editing every wrong line. The captions themselves are already AI — an ASR model guessed each word from sound. The correction is a second AI pass that knows things the transcriber did not: what your brand is called, which homophone you meant, where one sentence ends and the next begins. Done well, it turns "proofread every clip" into "spot-check a clip."
There are two places AI can intervene. Before transcription, you can bias the model toward the words it keeps getting wrong — a keyterm or custom-vocabulary list of your names, products, and jargon, plus a short prompt describing the audio. After transcription, you can run a language model over the raw transcript to repair the errors the ASR still made, using the surrounding context ASR ignores. The strongest setups do both, then keep a deterministic glossary as a backstop so the same word never slips twice.
This guide is the AI-first correction workflow: the pre-pass that prevents errors, the LLM post-pass that repairs them, and the automation that makes it run on every video without your attention. If you only want the fastest manual fix on a single platform, the [fix auto-caption errors](/how-to/fix-auto-caption-errors) guide covers the platform-by-platform hand edits; this page is about not doing them one at a time.
The steps
Know what AI correction can and cannot fix. ASR errors are systematic, not random, which is exactly why AI can target them. They cluster on proper nouns (your brand, products, guest names), homophones (their/there, to/too, its/it's), acronyms, numbers, and word boundaries where the model splits or joins the wrong syllables. A correction model with context resolves most of these because the sentence around the word tells it which spelling is right. What AI cannot recover is audio the transcriber never heard — a name buried under loud music, a mumbled word, crosstalk where two people speak at once. Fix the audio (duck the music, one speaker at a time) or those stay wrong no matter how good the corrector is.
Seed a keyterm / custom-vocabulary list before transcription. The cheapest correction is the error that never happens. Most managed ASR APIs let you pass a list of terms to bias the model toward: Deepgram calls it keyterm prompting, AssemblyAI calls it keyterms, others call it word boost or custom vocabulary. Load it with the words that break every time — your brand, product names, regular guests, industry jargon, acronyms — and the engine biases toward the correct spelling and capitalization on the first pass. Deepgram reports keyword recall on prompted terms rising toward 90%. Note that OpenAI's Whisper only exposes a 224-token `initial_prompt`, and OpenAI's own docs say it is not especially reliable for proper nouns, so on Whisper lean harder on the post-editing pass below.
Add a context prompt describing the audio. Separate from the term list, several ASR engines accept a short natural-language description of what the audio is about ("a marketing podcast about e-commerce, hosts Mara and Devon, discussing Shopify and Klaviyo"). This is the same idea as priming a language model: telling the transcriber the domain narrows its guesses toward the right vocabulary. It costs one sentence and measurably reduces the proper-noun and jargon errors that the correction pass would otherwise have to clean up. Write it once per show or brand and reuse it.
Run an LLM post-editing pass over the raw transcript. This is the core AI correction step. Feed the raw ASR transcript to a language model and ask it to correct transcription errors only — fix misheard words, homophones, punctuation, and casing, while preserving the exact wording and timing everywhere it is already right. Because the LLM sees the whole sentence, it resolves context-dependent errors that word-level tools cannot: it knows "we use Klaviyo for email" even when the ASR wrote "we use clay vo." Research on LLM-based ASR post-editing (judge-then-edit and rewrite approaches) shows the reliable gains come from keeping high-confidence spans untouched and only rewriting the uncertain ones — so instruct the model to be conservative and never paraphrase, or it will "improve" your wording and desync the captions.
Keep a deterministic glossary as the backstop. AI passes are probabilistic; a glossary is not. Maintain a short find-and-replace map of your known problem terms and their correct spellings, and apply it after the LLM pass. It guarantees the handful of words you care about most — the ones a customer would notice — are never wrong, even on the rare run where the model misses one. Think of it as a three-layer funnel: the keyterm list prevents most errors, the LLM pass repairs the contextual ones, and the glossary catches anything that survived both. The glossary is also where you encode decisions the AI cannot know, like whether it is "e-commerce" or "ecommerce" in your brand style.
Correct the shape, not just the words. Accurate text is only half of readable captions. Have the correction pass also re-segment: merge cues that break mid-phrase, split cues too long to read, and keep lines near 32 characters with no more than two on screen at once, so a phone viewer reads them at the ~17 characters per second most people manage. Some AI tools re-segment automatically from the corrected transcript. If a caption lands under the platform UI or on-screen chrome, nudge it into the readable middle third. A perfectly-spelled transcript that flashes two lines for half a second still fails the sound-off viewer.
Close the loop — proofread the high-stakes words, or caption from the script. AI correction shrinks the proofread; it does not delete it. Before anything customer-facing ships, eye-check the words that carry real cost if wrong: names, prices, claims, legal terms. This is a spot-check now, not a line-by-line rewrite. The deepest fix is to skip transcription entirely: when a video is generated from a written script, caption from those exact words rather than transcribing them back out of the audio, and the mis-hearing that starts the whole correction chain never occurs. That is the difference between correcting captions faster and not having to correct them.
Common gotchas
Do not let the LLM paraphrase. If the correction prompt is loose, the model rewrites for style, which both changes what you said and desyncs the captions from the audio. Instruct it to fix errors only and preserve wording and timing.
A keyterm list has a size limit — Deepgram caps keyterm prompting at 500 tokens per request, roughly 100 terms. Load it with your highest-value, most-often-wrong terms, not your entire product catalog.
Whisper's prompt parameter is weak for proper nouns by OpenAI's own admission. If you rely on Whisper, do not expect the pre-pass to carry the load — the LLM post-edit and glossary do the real work there.
AI correction cannot invent audio it never received. Music mixed loud under the voice is still the single biggest accuracy killer; duck it in the mix before transcribing, or the corrector is guessing too.
Auto-translated captions inherit every uncorrected error and add translation drift on top. Correct the source language first, then translate, and have a native speaker check anything customer-facing.
Burned-in captions are permanent pixels. Run the full correction pass and the proofread before you render, not after — once it is baked into the frames, fixing a word means re-exporting the whole clip.
Where Kompozy fits
The workflow above is real, and it is also a build project: an ASR provider with keyterm prompting wired in, an LLM prompted to post-edit conservatively, a glossary applied after it, and re-segmentation on top — assembled, maintained, and re-run on every clip. Kompozy is that stack already assembled, so the correction happens inside the render instead of in a pipeline you own.
For footage that has to be transcribed — [Clipped Shorts](/glossary/output-buckets) cut from long-form video, a podcast dropped in as a source — Kompozy runs a Whisper-class transcription and burns the captions in through an ffmpeg/libass step, already sized to the readable ~32-character line and positioned in the middle third. The part this page is about — biasing the transcription toward your vocabulary and repairing what it still gets wrong — is governed by the [Persona Brief](/glossary/persona-brief) that steers every render: it holds your brand, product, and recurring names steady, which is the keyterm list and the correction glossary without you maintaining either. The LLM that governs the copy is the same class of model the post-editing step calls for, so context-dependent fixes happen in the same pass that produces the text.
For the video Kompozy generates from a written script — [Persona Shorts](/glossary/persona-shorts) and Persona HeyGen — it takes the deepest fix on this page by default: the captions render from the known script, never transcribed back out of the avatar's speech, so the proper-noun and homophone errors that trigger the whole correction chain do not occur. Either way, a per-post review pipeline is where the one remaining spot-check lives before [Autopilot](/glossary/autopilot) fans the captioned video across the eight social platforms plus blog and email on schedule. Creator ($49/mo for 2,500 credits) fits a solo creator done hand-correcting the same names every week; Pro ($299/mo for 18,000 credits) suits high-volume, multi-brand output where per-clip correction never scales; Enterprise is custom.
Frequently asked questions
What is AI-powered auto-caption correction?
It is using AI to fix the errors in automatic captions, rather than editing them by hand. Two layers do the work: before transcription, a keyterm or custom-vocabulary list biases the ASR model toward your names and jargon; after transcription, a language model reads the raw transcript and repairs misheard words, homophones, and punctuation using the surrounding context. A deterministic glossary usually backstops both.
Can an LLM fix speech-recognition errors?
Yes, and it is a well-studied technique. Because a language model sees the whole sentence, it resolves context-dependent errors that the transcriber missed — it can tell that "clay vo" should be "Klaviyo" from the sentence around it. The reliable approach keeps high-confidence text untouched and only rewrites uncertain spans, so the model must be told to correct errors only and never paraphrase, or it will change your wording and break caption timing.
How do I stop auto-captions misspelling brand and product names?
Seed a keyterm or custom-vocabulary list (Deepgram keyterm prompting, AssemblyAI keyterms, and similar features on other ASR APIs) with your brand, product, and people names so the engine biases toward the right spelling on the first pass. Keep a find-and-replace glossary of those same terms as a backstop after the AI pass. Together they get proper nouns close to reliable; captioning from a known script is the only complete fix.
Does AI caption correction remove the need to proofread?
No — it shrinks the proofread to a spot-check. AI passes are probabilistic and will occasionally miss a term, so before anything customer-facing ships, eye-check the high-stakes words: names, prices, claims, legal terms. That is far faster than reading every line, but skipping it entirely on important content is how a wrong price or a misspelled partner name reaches an audience.
Is it better to correct captions with AI or caption from a script?
Captioning from a script is strictly better when you have one, because there is nothing to mis-hear — the errors never occur. AI correction is for the footage you did not script: interviews, podcasts, talks, screen recordings. Use script-based captions for generated or read-from-teleprompter video, and the AI correction stack for everything transcribed from real audio.
What tools do the AI correction steps?
The pre-pass runs on your ASR provider: Deepgram, AssemblyAI, and others expose keyterm/custom-vocabulary and context prompting; Whisper offers a weaker 224-token prompt. The post-pass is any capable LLM (Claude or GPT-class) prompted to correct errors conservatively. The glossary is a simple find-and-replace. A content engine that chains all three inside the render removes the need to assemble the stack yourself.