// GUIDE · 2026-09-09

Performance-preserving dubbing: how lip-matched AI localization keeps the on-camera performance (2026)

For most of dubbing's history the deal was a bad trade: you got the words in your language, but the actor's mouth kept moving in the original one, and every close-up broke the illusion. In 2026 that trade is being renegotiated. A new class of localization — call it lip-matched, visual, or performance-preserving dubbing — reshapes the visible mouth to fit the translated audio, so the face and the dialogue finally agree. Amazon put a headline on it when Prime Video shipped AI-and-VFX lip-sync on the series Maxton Hall in September 2026, and drew the line precisely: human voice actors still perform the dub, and the synthetic layer only adjusts the lips, framed as preserving the integrity of the original artistic vision. That single detail — real voice, synthetic mouth — is the key to understanding a field that most people flatten into 'AI dubbing.' This guide takes apart what 'preserving the performance' actually means, which parts of a performance each localization approach keeps and which it quietly discards, why the studio-only version of this became a routine content decision, where matched-mouth dubbing still visibly breaks, and the strategic fork every creator now faces: retrofit a finished clip, or generate the performance in each language from the start.

Last verified · 2026-09-09 · by Moe Ameen

The trade dubbing always asked you to make

For most of its history, dubbing asked for a compromise you could not avoid: you got the dialogue in your language, but the actor's mouth kept moving in the original one. On a wide shot or a voiceover it did not matter. On a close-up it broke the illusion every time — the lips finishing a sentence the audio had already ended, or a plosive landing on a closed mouth. That mismatch is the single most recognizable tell of a dub, and it is why so many viewers instinctively prefer subtitles for anything with a face in frame. The words were localized; the performance was not.

What changed in 2026 is that the mouth became editable. A new class of localization — variously called lip-matched, visual, or performance-preserving dubbing — reshapes the visible mouth to fit the translated audio, so the face and the dialogue finally agree. This is a different job from the three most people mean by 'AI video translation,' and the full breakdown of subtitles versus dubbing versus lip-sync is worth reading alongside this. Here the focus is narrower and more specific: what it actually means to preserve a performance through localization, which approaches keep it and which quietly discard it, and what that means for anyone publishing video to more than one language.

Prime Video made it a headline — and drew the line precisely

The clearest signal that this crossed from research into product came from Amazon. On September 9, 2026, Prime Video announced a lip-sync capability that adjusts actors' on-screen mouth movements so they line up with dubbed dialogue, powered by a combination of AI and VFX. It debuted on the German romance series Maxton Hall, available globally in English on Seasons 1 and 2, with the third season built to include it. The stated goal was exactly the tell described above: fix the mouth that keeps moving after the line is done, so global content feels native rather than overdubbed. The full announcement and what it means for creators sits in the news log.

One detail defines what this is and is not: the dub itself is still human. Professional voice actors record the translated dialogue, and the AI-and-VFX layer only reshapes the visible mouth to match that human performance. Amazon framed the whole thing as running under creative oversight — VP of technology Raf Soltanovich described the aim as a more seamless, immersive way to enjoy global content, done so that 'the integrity of their artistic vision is preserved.' That is not marketing gloss; it is the actual design decision. The voice performance stays real, the physical performance stays real, and only the lips — the one part that has to change to match new audio — are synthetic.

The two switches that define every localized talking head

Almost everyone flattens this space into 'AI dubbing,' which hides the distinction that actually governs quality. A cleaner model is two independent switches. Switch one: is the voice the viewer hears real (a human actor, in the source or the target language) or synthetic (a cloned or generated voice)? Switch two: is the mouth on screen real (untouched original footage) or synthetic (re-animated to match the audio)? Every localization method is one of the four combinations, and each preserves a different slice of the original performance.

Voice-only dubbing — keeps the face, breaks the mouth

The traditional approach: replace the audio with a translated voice, leave the video untouched. The full physical performance survives — expression, eyes, gesture, head motion — but the mouth still moves in the source language, so a talking-head close-up gives it away. With voice cloning the translated voice can still sound like the original speaker, and good tools keep the music and ambience. This is the right call when the mouth is not the focus: a voiceover over B-roll, a wide shot, a screen recording. Nobody is watching the lips, so there is nothing to break.

Lip-matched (visual) dubbing — keeps the face, fixes the mouth

Prime Video's approach and the direction the flagship tools are moving: keep the real footage, re-animate only the mouth region so the lips form the translated words. This is the only method that fully solves close-up talking-head footage while leaving the rest of the performance intact — the actor's eyes, brow, and gestures are still theirs, and now the mouth agrees with the audio. It is also the most expensive and the most visible when it fails, because a wrong mouth reads as uncanny faster than a wrong voice. The voice underneath it can be human (Prime Video) or synthetic (most creator tools); the visual step is the same either way. The concept behind it is what AI lip sync describes, applied to real footage rather than a generated avatar.

Generate-in-language — the mouth was never wrong

The fourth combination sidesteps the retrofit entirely: generate the video natively in the target language from a translated script, so the mouth is correct from the first frame because it was never speaking the source language. There is no re-animation to get wrong and no drift to review. The trade is that you are not preserving a specific filmed performance — you are substituting a new one, typically an avatar or persona, that is consistent by construction. For a real actor's flagship scene that is the wrong tool; for recurring content where the point is a consistent on-brand presenter rather than a particular take, it is often the right one.

What 'performance' actually means — and which parts survive

It helps to be precise about what a performance is, because 'preserving' it is doing a lot of work in the marketing. A performance is two layers: a vocal layer (emotion, prosody, timing, the specific way a line is delivered) and a physical layer (facial expression, eye movement, gesture, the mouth). Different localization methods keep different layers. Voice-only dubbing with a strong target-language actor preserves the whole physical layer but replaces the vocal layer with a new performer's — usually a fair trade, since a native-language delivery is the point. Prime Video's human-voice, synthetic-mouth approach preserves the physical layer minus the mouth and hands the vocal layer to a human dub actor, which is why it can claim to protect the artistic vision: a person, not a model, is still performing every word.

A full synthetic-voice dub is where the most is lost: a cloned voice reproduces timbre but often flattens the emotional dynamics and comic timing that made the original land, and on less-resourced languages the drop is sharper. Generate-in-language does not preserve the source performance at all; it manufactures a new, consistent one. None of these is strictly better — the right question is which layer of the performance actually matters for your footage. A dramatic scene lives on vocal nuance and a real mouth; a weekly explainer lives on a consistent presenter and correct information, where a generated performance is not a loss at all.

Where lip-matched dubbing still visibly breaks

The honest picture: matched-mouth dubbing is very good on clean, front-facing, single-speaker footage and genuinely unreliable outside those conditions. Profile and three-quarter angles hide part of the mouth and confuse the re-animation. Fast or overlapping speech breaks both the transcription that feeds the dub and the sync that matches the lips. Multiple speakers require accurate diarization, and a line assigned to the wrong face falls apart. Heavy head motion and close-ups shot at an angle degrade the mouth region specifically. And because a wrong mouth is more jarring than a wrong voice, a bad lip-match is worse than no lip-match — the uncanny mouth actively repels the viewer the clean overdub only mildly annoyed.

Length is the quieter failure. Voiced dialogue runs to different durations in different languages — German and Spanish dubs frequently overrun their English source — so the lips are being asked to sync to audio the shot cannot physically hold, and the result either rushes or spills past the cut. This is why the reliable pattern is still AI for the first pass at volume with a native-speaker review on anything the brand rides on, exactly as covered in how to localize video for multiple languages and, for the mouth step specifically, how to lip-sync a dubbed video. Treat the raw output as a strong draft, never a finished master.

Why the studio-only version became routine

Performance-preserving dubbing is not new as a capability — high-end studios have re-animated mouths for years. What changed is the price. Studio dubbing with lip-sync historically ran roughly $80 to $200 or more per finished minute and took days to weeks per language. AI dubbing in 2026 lands at entry tiers around $0.24 to $2.40 per minute for lip-synced output, dropping toward $0.10 for audio-only — a 90 to 98 percent cut on the production line, with turnaround compressed by as much as 92 percent. The honest asterisk: once native-speaker review, rework, and compliance are added, most teams land at two to four times the tool's headline rate. That is still an order of magnitude below a studio quote, but it is not free.

The consequence is a strategy change, not just a cheaper invoice. When a language version cost a thousand dollars and a week, you localized your single most important asset into your two biggest markets. When it costs a few dollars and minutes, localizing a whole library into a dozen languages becomes a routine content decision, and the bottleneck moves from producing the translation to reviewing it and getting every version distributed. The same economics drive the creator-side dubbing playbook: the translation stopped being the hard part; the hard part is quality control and reach.

The walled-garden gap this leaves creators

Here is the catch for anyone who is not a streaming service. Prime Video's polish lives inside the Prime Video player, on Amazon's own catalog. A creator publishing to YouTube, TikTok, Instagram, and LinkedIn gets none of it and has to build localized video themselves. Platform-native dubbing closes part of the gap — YouTube's free auto-dubbing generates switchable translated tracks, and Instagram and Facebook dub Reels in your own voice with lip adjustment — but each is per-platform, tied to a fixed language set, and offers little control over the visual match. The bar performance-preserving dubbing sets is global; the tooling to hit it is fragmented across a dozen products and locked catalogs.

So the practical field for a creator is: native platform features first (free, decent, limited), a dedicated dubber like HeyGen or ElevenLabs when you need a specific language or higher fidelity, and a native-speaker review before anything customer-facing ships. The roundup of AI video translators breaks those tools down by language coverage, voice cloning, and how lip-sync is metered. What none of them solve is the part after the dub — turning one localized asset into the posts, clips, and companion content each market actually needs to find you.

How to choose for your footage

Work backward from the shot and the stakes. If the value is information and the audience will read, subtitles are cheapest and safest — the real voice survives untouched. If the speaker is off-frame or incidental, voice-only dubbing gives a native-language track without paying the lip-sync premium. If a real person is talking to camera and the market will judge the video on how native it feels, lip-matched dubbing is the only method that fully solves it, and it earns its cost on your highest-value assets. And if this is recurring content rather than a one-off — a weekly series, an ongoing presence in several markets — generating the video in each language from the start beats retrofitting every clip, because you never pay the re-animation tax or review the drift.

That last fork is the one that decides whether localization is a project or a capability. Dubbing a finished video is a translate-then-distribute motion: it solves the clip. Generating in-language is a produce-in-parallel motion: it solves the calendar. Most serious multi-market operations end up running both — lip-matched dubbing on the hero assets where a specific performance matters, and native generation for the steady cadence where consistency matters more than any single take.

Where Kompozy fits: the performance you keep is the brand

Kompozy sits on the generate-in-language side of that fork, and it reframes what 'preserving the performance' means. It is an AI content generation and multi-platform publishing engine, not a dubber — you will not paste a finished clip of a real actor into it and get a lip-matched version back. What it preserves across markets is not one filmed take but a consistent brand identity: a face-locked AI Influencer persona whose voice and mouth are generated together in the target language from a translated script, through HeyGen's native text-to-speech. Because the audio and the lips are produced in the same step, there is nothing to retrofit and no drift to review — the mouth was never wrong. The performance that stays constant from English to Spanish to Portuguese is your persona's voice and look, held to one Persona Brief so nothing goes off-message market to market.

That is a different guarantee than a studio's. Prime Video preserves a specific actor's artistic vision on a specific title; a Persona Short or Persona HeyGen video preserves your recurring on-brand presenter across every language you publish in, on a cadence. And Kompozy then does the part the streaming feature deliberately is not for — distribution. One localized talking-head asset fans into captioned vertical clips, a brand-exact carousel, quote graphics, a blog article, and an email newsletter in that market's language, and Autopilot schedules the batch across eight social platforms plus blog and email behind a per-post review gate where a native speaker can sharpen a line before it ships.

Two honest boundaries keep this credible. If your job is to localize an existing video of a real person and keep that person's performance, a dedicated visual-dubber — HeyGen for lip-sync, ElevenLabs for voice, and the concept behind AI dubbing generally — is the tool, and Kompozy is not competing with it. And native generation is not a magic quality button: volume without a review gate produces localized filler that audiences and platforms now demote. Kompozy earns its place when localization is not a one-time translation task but a standing requirement to produce and publish on-brand content in several languages, every week, without a separate production run per market — the calendar problem, not the clip problem.

Frequently asked questions

What is performance-preserving (lip-matched) dubbing?

It is video localization that reshapes the speaker's on-screen mouth to match the translated audio, instead of only swapping the voice. Standard dubbing replaces the audio but leaves the lips moving in the original language, so a close-up gives away the dub. Performance-preserving dubbing adds a video step — often AI plus VFX — that re-animates the mouth to form the new words, so the face and the dialogue agree and the actor's expression, timing, and gestures survive intact. Amazon's Prime Video shipped a version of this on the series Maxton Hall in 2026, keeping human voice actors for the audio and using the synthetic layer only for the lips.

Does lip-matched dubbing use an AI voice or a human voice?

It depends on the approach, and the difference matters. Prime Video's implementation keeps the dub human-performed — professional voice actors record the translated dialogue — and uses AI and VFX only to adjust the visible mouth. Many creator-facing tools do the reverse or both: they generate a cloned or synthetic voice and re-animate the mouth to match it. So 'lip-matched dubbing' describes the visual step (the mouth), not the audio source. The cleanest mental model is two independent switches — is the voice real or synthetic, and is the mouth real or synthetic — and the four combinations preserve very different amounts of the original performance.

Why does preserving the performance matter for dubbing?

Because a performance is more than its words. The old tell of dubbing — a mouth that keeps moving after the translated line ends, or stops before it — pulls a viewer out of the story and reads as low-effort. Matching the mouth to the audio removes that tell, so the actor's facial expression, eye movement, timing, and body language carry through instead of fighting a mismatched mouth. As a top streamer ships this on a flagship title, audiences start reading the old mismatched overdub as cheap, which raises the localization bar for everyone publishing video, not just studios.

Where does lip-matched dubbing still break?

On the hard footage: profile and three-quarter angles where the mouth is partly hidden, fast or overlapping speech that confuses both transcription and sync, multiple speakers that require accurate diarization, heavy head motion, and close-ups shot at an angle. Voiced length also differs by language — German and Spanish dubs often run longer than English and overrun the scene — so timing has to be managed or the lips sync to audio the shot cannot hold. And a mismatched or rubbery mouth reads as uncanny far faster than an audio glitch, so bad lip-sync is more damaging than no lip-sync.

How much does AI lip-sync dubbing cost versus a studio?

Studio dubbing with lip-sync historically ran roughly $80–$200+ per finished minute and took days to weeks. AI dubbing in 2026 lands at entry tiers around $0.24–$2.40 per minute for lip-synced output, dropping toward $0.10 for audio-only, a 90–98% cut on the production line and up to a 92% reduction in turnaround. The honest caveat: once you add native-speaker review, rework, and compliance, most teams land at two-to-four times the tool's headline rate — still an order of magnitude below a studio quote, but not the free number the marketing implies.

Should I dub an existing video or generate it in each language?

For a back-catalog or a one-off flagship asset, dub the finished video — you are reusing work you already made, and lip-matched dubbing is what makes a real presenter look native. For recurring content in several markets, generating the video natively in each language from a translated script is usually the better path: the mouth is correct from the first frame because it was never speaking the source language, so there is no re-animation to get wrong and nothing to drift. Many serious multi-market operations run both — dub the hero assets, generate the ongoing cadence.

The direct answer

Performance-preserving (lip-matched) dubbing localizes video by reshaping the speaker's on-screen mouth to match the translated audio, not just swapping the voice — so the actor's face, expression, and timing survive instead of fighting a mismatched mouth. Prime Video shipped a version on Maxton Hall in September 2026, keeping human voice actors and using AI plus VFX only for the lips. The approach removes the classic dubbing tell but still breaks on profile shots, multiple speakers, and fast speech, and for recurring content, generating video natively in each language often beats retrofitting a finished clip.

Get started → · ← All guides · Compare Kompozy vs other tools