Dubbed a video and the mouth no longer matches? Here's how to lip-match on-camera footage to the new language, preserve the performance, and avoid the tells.
Last verified · 2026-09-09 · by Moe Ameen
Standard dubbing swaps the audio and leaves the picture alone, so the speaker's mouth keeps moving in the original language — fine on a wide shot or a voiceover, a giveaway on any close-up. Lip-matching (also called visual or performance-preserving dubbing) adds the step that fixes it: it re-animates only the mouth region of your real footage so the lips form the translated words, while the rest of the performance — the eyes, the expression, the head motion — stays exactly as filmed. This is the step that makes a localized talking-head video look native instead of overdubbed, and it is the same technique a major streamer like Prime Video now ships on flagship titles.
This walkthrough is specifically about that visual step for real on-camera footage of a person, not about generating a talking video from a photo (that is a different task) or about the audio side of translation. The goal is a dubbed clip where the mouth agrees with the new audio and the speaker still looks like themselves. For the broader picture of subtitles-vs-dubbing-vs-lip-sync and when each is worth it, read the [performance-preserving dubbing guide](/guides/performance-preserving-dubbing) alongside this; for the full four-step translate-and-dub pipeline, the [AI video translation how-to](/how-to/translate-a-video-with-ai) goes deep on the audio half.
Lip-matching your own footage into another language is fine. Re-animating a real person's mouth to say words they never spoke can create a deepfake — only lip-match footage of yourself, of people who have given clear consent, or of licensed talent, and never to put false words in someone's mouth. If you clone a voice for the dub, clone only one you own or are licensed to use. AI-altered video is subject to platform disclosure rules (TikTok, YouTube, Instagram and others) and increasingly carries provenance signals such as SynthID or C2PA — label AI-modified media per each platform's policy, and check whether your target markets require it.
Lip-matching a real presenter is per-clip labor that never ends: every new video is a fresh translate, re-sync, inspect, and fix cycle on footage you already shot, and the sync can still drift on the shots that matter most. For a single hero asset that is worth it, and a dedicated visual-dubber like HeyGen or sync. does exactly that job — Kompozy is not a re-sync tool for arbitrary footage, and you should reach for those when the task is to localize one specific clip of a real person.
Where [Kompozy](/) changes the math is the ongoing channel, by removing the re-sync step entirely. It is an AI content generation and multi-platform publishing engine, and its [Persona Shorts](/glossary/persona-shorts) and Persona HeyGen formats generate a talking-head avatar that speaks the target language natively through HeyGen's text-to-speech — the voice and the mouth are produced in the same render, so there is no mouth to retrofit and no drift to hunt for. The per-clip inspect-and-fix loop this tutorial walks through becomes a non-issue for recurring content, because the lips were never out of sync. A [Persona Brief](/glossary/persona-brief) holds your voice and look constant across every language, so the presenter is recognizably you from English to Spanish to Portuguese.
The review gate is where the native-speaker check from step eight becomes structural instead of something you remember to do: a person approves each localized piece — sharpening a headline, catching a translation miss — before [Autopilot](/glossary/autopilot) schedules and publishes it across eight social platforms plus blog and email, and the same source fans into captioned clips, a brand-exact carousel, quote graphics, a blog article, and a newsletter in that market's language, so a localized launch is a full footprint rather than one dubbed clip. Creator ($49/mo for 2,500 credits) fits a solo operator testing a second-language channel; Pro ($299/mo for 18,000 credits) carries a serious multi-market cadence with the companion formats per launch; Enterprise is custom for full localization programs.
Dubbing replaces the audio with a translated voice but leaves the video untouched, so the speaker's mouth still moves in the original language. Lip-matching (visual dubbing) adds a step that re-animates the on-screen mouth to form the translated words, so the lips agree with the new audio. Lip-matching is what makes a close-up talking-head video look native instead of overdubbed; dubbing alone is fine when the mouth is off-frame or incidental.
Done well, only the mouth region. The value of lip-matching over regenerating the video is that the original expression, eyes, and head motion are preserved — only the lips change to match the new audio. If a tool visibly flattens or freezes the rest of the face, the result looks pasted-on; that is a sign to switch tools or shorten the segment.
HeyGen's video translation and lip-sync tools and sync. (sync-3) are built to re-sync real footage, re-rendering the mouth to match a new audio track. Some all-in-one dubbers run translation, voice cloning, and lip-sync in one pass for speed at the cost of stage-by-stage control. Photo-to-talking-avatar generators are a different category and are not the right tool for localizing existing footage of a real person.
Usually one of four things: the mouth region is blurry or the jaw boundary shifts, the lips lead or lag the audio, the dub overran the shot so the sync is chasing audio the frame cannot hold, or the surrounding face was flattened. Test a short segment, fix the length mismatch first, use a diffusion-based tool for sharper teeth and jaw, and confirm the eyes and expression are still the original take.
For a one-off asset of a real person whose specific performance matters, lip-match the existing footage. For recurring content across several markets, generating the video natively in each language from a translated script is usually better — the mouth is correct from the first frame, so there is no re-animation to get wrong and no drift to review. Many multi-market operations lip-match hero assets and generate the ongoing cadence.