// HOW-TO · LOCALIZATION

How to lip-sync a dubbed video so the mouth matches the new language (2026)

Dubbed a video and the mouth no longer matches? Here's how to lip-match on-camera footage to the new language, preserve the performance, and avoid the tells.

Last verified · 2026-09-09 · by Moe Ameen

Standard dubbing swaps the audio and leaves the picture alone, so the speaker's mouth keeps moving in the original language — fine on a wide shot or a voiceover, a giveaway on any close-up. Lip-matching (also called visual or performance-preserving dubbing) adds the step that fixes it: it re-animates only the mouth region of your real footage so the lips form the translated words, while the rest of the performance — the eyes, the expression, the head motion — stays exactly as filmed. This is the step that makes a localized talking-head video look native instead of overdubbed, and it is the same technique a major streamer like Prime Video now ships on flagship titles.

This walkthrough is specifically about that visual step for real on-camera footage of a person, not about generating a talking video from a photo (that is a different task) or about the audio side of translation. The goal is a dubbed clip where the mouth agrees with the new audio and the speaker still looks like themselves. For the broader picture of subtitles-vs-dubbing-vs-lip-sync and when each is worth it, read the [performance-preserving dubbing guide](/guides/performance-preserving-dubbing) alongside this; for the full four-step translate-and-dub pipeline, the [AI video translation how-to](/how-to/translate-a-video-with-ai) goes deep on the audio half.

The steps

  1. Get the dub right before you touch the mouth. Lip-matching syncs the lips to the audio, so a wrong or badly-timed dub produces a wrong mouth. Finish the audio side first: translate the script (with a native-speaker or glossary pass on proper nouns and idioms), generate or record the target-language voice, and — ideally — clone the original speaker's voice so the dub still sounds like them and keep the background music and ambience intact. Only take a clean, final dub into the visual step. Fixing the translation after you have re-animated the mouth means re-rendering everything.
  2. Start from clean, front-facing source footage. The mouth re-animation is capped by how well the tool can see the mouth. Front-facing, evenly lit footage with the mouth unobstructed re-syncs cleanly; hard profiles, hands or mics across the jaw, heavy shadow, and fast head motion all degrade it. If you are still filming, shoot the talking-head segments straight-on knowing they will be localized. If you are working with existing footage, note which shots are clean enough to lip-match and which you will have to handle another way (voice-only over B-roll, or a cutaway).
  3. Pick a visual-dubbing tool that re-renders only the mouth. You want a tool that changes the mouth region and leaves the rest of the frame untouched, so the original performance survives. HeyGen's video translation and lip-sync tools and sync. (its sync-3 model) are built for re-syncing real footage; some all-in-one dubbers run translate, voice, and lip-sync in a single pass, which is faster but gives you less control over each stage. Match the tool to your footage — mouth-only re-sync for a real presenter, not a photo-to-avatar generator, which is a different job covered in [how to make AI lip sync videos](/how-to/make-ai-lip-sync-videos).
  4. Match the dub length to the shot, not the other way around. Voiced dialogue runs to different lengths in different languages — German and Spanish dubs frequently overrun their English source — and if the audio is longer than the scene, the lips are being asked to sync to words the shot cannot physically hold. Before rendering, check each segment's dubbed audio against the original clip length. Re-pace or lightly time-stretch the dub, tighten the translation, or add a beat of slack in the edit so the mouth is syncing to audio that fits the frame.
  5. Render the lip-match and inspect the mouth at full size. Run one short segment first, then watch the mouth region zoomed in, not the thumbnail. Hunt for the specific artifacts: a blurry or shifting jaw boundary, teeth that smear or flicker, lips that lead or lag the audio, and — on longer clips — sync that drifts as the segment runs. Catching a mispronounced brand name or a wobbly mouth on one 10-second test is far cheaper than on the whole video. Do not approve a render just because the first two seconds line up.
  6. Confirm the rest of the performance survived. The whole point of lip-matching over a full re-generation is that only the mouth should change. Check that the eyes, brow, expression, and head motion are still the original take and that the tool has not flattened or frozen the face around the mouth. If the surrounding performance looks deadened or the mouth region is visibly pasted on, switch tools or shorten the segment — a technically-synced mouth on a lifeless face is worse than a clean voice-only dub.
  7. Handle multi-speaker segments and bad angles separately. Multiple speakers require the tool to assign each line to the right face; a misassignment falls apart, so segment the clip by speaker and lip-match each pass cleanly. For shots the tool cannot handle — profiles, obscured mouths, crowd scenes — do not force a bad lip-match. Leave those as voice-only dub (the mouth is not the focus there anyway) or cut around them. A mix of clean lip-matched close-ups and voice-only wide shots reads far better than a uniformly forced sync.
  8. Localize the packaging, review with a native speaker, then publish. A lip-matched clip under an original-language title, thumbnail, and captions is a localized video inside an un-localized listing. Translate the title and description, recreate on-screen text and thumbnails in the target language, and regenerate captions in-language. Then have a native speaker watch the full cut for translation misses and any residual sync drift before it ships — this is the review that separates a first pass from a finished master. Publish per market, reframing per platform rather than dropping one file everywhere.

Common gotchas

  • Profiles and obscured mouths break lip-matching. The tool needs a clear, front-facing view of the mouth — hard angles, hands, mics, and shadow across the jaw all degrade or ruin the re-sync.
  • A wrong mouth is worse than no lip-sync. A rubbery or mistimed mouth reads as uncanny far faster than a clean overdub reads as dubbed — if a segment will not sync well, fall back to voice-only.
  • Dub length differs by language and overruns the shot. German and Spanish often run longer than English; sync the audio to the scene length first, or the lips chase words the frame cannot hold.
  • Sync holds on short clips and drifts on long ones. Many tools stay locked for 30–60 seconds and slip on multi-minute footage — test a long sample or render in shorter segments.
  • Only the mouth should change. If the tool freezes or flattens the eyes and expression around a synced mouth, you have lost the performance you were trying to preserve — switch tools.
  • On-screen text is never dubbed. Lower-thirds, slide text, and baked-in graphics stay in the original language — recreate them per market or design without on-screen text.
  • Multiple overlapping speakers confuse line assignment. Segment by speaker and lip-match each pass; a whole-clip pass mismatches lines to faces.
Legal note

Lip-matching your own footage into another language is fine. Re-animating a real person's mouth to say words they never spoke can create a deepfake — only lip-match footage of yourself, of people who have given clear consent, or of licensed talent, and never to put false words in someone's mouth. If you clone a voice for the dub, clone only one you own or are licensed to use. AI-altered video is subject to platform disclosure rules (TikTok, YouTube, Instagram and others) and increasingly carries provenance signals such as SynthID or C2PA — label AI-modified media per each platform's policy, and check whether your target markets require it.

Where Kompozy fits

Lip-matching a real presenter is per-clip labor that never ends: every new video is a fresh translate, re-sync, inspect, and fix cycle on footage you already shot, and the sync can still drift on the shots that matter most. For a single hero asset that is worth it, and a dedicated visual-dubber like HeyGen or sync. does exactly that job — Kompozy is not a re-sync tool for arbitrary footage, and you should reach for those when the task is to localize one specific clip of a real person.

Where [Kompozy](/) changes the math is the ongoing channel, by removing the re-sync step entirely. It is an AI content generation and multi-platform publishing engine, and its [Persona Shorts](/glossary/persona-shorts) and Persona HeyGen formats generate a talking-head avatar that speaks the target language natively through HeyGen's text-to-speech — the voice and the mouth are produced in the same render, so there is no mouth to retrofit and no drift to hunt for. The per-clip inspect-and-fix loop this tutorial walks through becomes a non-issue for recurring content, because the lips were never out of sync. A [Persona Brief](/glossary/persona-brief) holds your voice and look constant across every language, so the presenter is recognizably you from English to Spanish to Portuguese.

The review gate is where the native-speaker check from step eight becomes structural instead of something you remember to do: a person approves each localized piece — sharpening a headline, catching a translation miss — before [Autopilot](/glossary/autopilot) schedules and publishes it across eight social platforms plus blog and email, and the same source fans into captioned clips, a brand-exact carousel, quote graphics, a blog article, and a newsletter in that market's language, so a localized launch is a full footprint rather than one dubbed clip. Creator ($49/mo for 2,500 credits) fits a solo operator testing a second-language channel; Pro ($299/mo for 18,000 credits) carries a serious multi-market cadence with the companion formats per launch; Enterprise is custom for full localization programs.

Frequently asked questions

What is the difference between dubbing and lip-matching a video?

Dubbing replaces the audio with a translated voice but leaves the video untouched, so the speaker's mouth still moves in the original language. Lip-matching (visual dubbing) adds a step that re-animates the on-screen mouth to form the translated words, so the lips agree with the new audio. Lip-matching is what makes a close-up talking-head video look native instead of overdubbed; dubbing alone is fine when the mouth is off-frame or incidental.

Does lip-matching change the whole face or just the mouth?

Done well, only the mouth region. The value of lip-matching over regenerating the video is that the original expression, eyes, and head motion are preserved — only the lips change to match the new audio. If a tool visibly flattens or freezes the rest of the face, the result looks pasted-on; that is a sign to switch tools or shorten the segment.

Which tools lip-sync real footage into another language?

HeyGen's video translation and lip-sync tools and sync. (sync-3) are built to re-sync real footage, re-rendering the mouth to match a new audio track. Some all-in-one dubbers run translation, voice cloning, and lip-sync in one pass for speed at the cost of stage-by-stage control. Photo-to-talking-avatar generators are a different category and are not the right tool for localizing existing footage of a real person.

Why does my lip-synced dub look fake?

Usually one of four things: the mouth region is blurry or the jaw boundary shifts, the lips lead or lag the audio, the dub overran the shot so the sync is chasing audio the frame cannot hold, or the surrounding face was flattened. Test a short segment, fix the length mismatch first, use a diffusion-based tool for sharper teeth and jaw, and confirm the eyes and expression are still the original take.

Is it better to lip-match a video or regenerate it in the new language?

For a one-off asset of a real person whose specific performance matters, lip-match the existing footage. For recurring content across several markets, generating the video natively in each language from a translated script is usually better — the mouth is correct from the first frame, so there is no re-animation to get wrong and no drift to review. Many multi-market operations lip-match hero assets and generate the ongoing cadence.

Related tutorials

← All how-to guides · Get Started