// AI NEWS · FEATURE

Descript Expands AI Dubbing With Lip-Synced Video, Regenerating the Speaker's Mouth to Match 30+ Languages

Descript's translation-and-dubbing now goes beyond swapping the audio: it uses generative AI to regenerate the lower half of the speaker's face so the mouth moves to the dubbed language, and it can translate on-screen text layers too — putting believable localized video within reach of a solo creator.

2026-08-17 · by Moe Ameen

What happened

Descript, the transcript-based AI audio and video editor, has expanded its translation and dubbing so a localized video can look as native as it sounds. The tool already dubbed a recording into more than 30 languages using native-sounding AI voices generated from the transcript. The expansion adds AI lip sync: after it produces the translated voiceover, Descript automatically regenerates the lower half of the speaker's face to match the new language's sounds, pacing, and mouth movements, so the video no longer reads as a dub laid over a mismatched mouth.

The lip sync is generative rather than rotoscoped. Descript encodes the original video and reference frames of the speaker's face into a learned latent space, generates new mouth movements driven by the translated speech, and blends them back over the untouched background and upper face — keeping the speaker's identity, lighting, and teeth consistent. It also translates text layers, so on-screen graphics and captions can be localized alongside the audio instead of being left in the source language.

Availability follows Descript's plan tiers. Dubbing and lip sync run on the paid Creator plan and above, and the feature consumes AI Credits; longer videos take longer to render. On the Creator plan the translation is automatic, while Business and Enterprise plans let you fine-tune the translation directly in the transcript before it renders and apply a "Do Not Translate" list from Brand Studio so product names and brand terms survive localization unchanged. Descript has documented the workflow in its help center and detailed the engineering behind multilingual dubbing at scale in a case study with OpenAI.

The framing to keep straight: this is localization of an existing recording — one polished video turned into many language versions with matched mouths. It is not, on its own, a system for cutting those versions into per-platform posts or distributing them across every network in each region. Confirm the current language list, plan gates, and credit costs on descript.com, since AI features and their limits change often.

Why it matters for creators

  • Lip-synced dubbing removes the biggest tell of a machine translation. A voiceover over an unmoving or mismatched mouth signals "dubbed" instantly; regenerating the mouth to the new language makes a localized talking-head video watchable to a native audience.
  • It puts studio-grade localization on a solo budget. Matched-mouth dubbing used to mean a voice actor and a VFX pass; a creator can now generate a French, German, or Spanish cut of the same video from the transcript in the same app.
  • Translating text layers matters more than it sounds. A dubbed video with English lower-thirds still reads as foreign; localizing on-screen graphics and captions is what makes the whole frame feel native, not just the audio.
  • The brand-safety controls decide whether it is usable at work. Transcript-level correction and a Do Not Translate list are what separate a rough auto-dub from a version a company will actually publish — and those sit on the higher tiers.
  • Localization multiplies your output, then hands you a distribution problem. One video becomes ten language masters — and each still has to be clipped, captioned, and posted to the right regional feeds, which is a separate, larger job than making the dubs.

How to act on this with Kompozy

Descript's lip-synced dubbing solves localization: one recording becomes a stack of language masters with matched mouths. It does not solve what comes after — turning each of those masters into per-platform, per-region posts and getting them out everywhere. That is the gap [Kompozy](/) is built for. Take a Descript master — the source cut and each dubbed version — into Kompozy, and it fans every one into the feed-native set: vertical [Clipped Shorts](/glossary/clipped-short) with word-synced captions in the matching language, a [Carousel](/glossary/hyperframes) and [Quote Graphics](/glossary/output-buckets) of the key points, a [Blog Article](/glossary/output-buckets) and an [Email Newsletter](/glossary/output-buckets) written from the transcript, each held to one voice by the [Persona Brief](/glossary/persona-brief).

The move a creator makes this week: localize the video once in Descript, then let Kompozy handle the multiplication. It reframes each language cut to 9:16, 1:1, and 16:9, generates the surrounding formats per version, and schedules and publishes across the eight social platforms plus blog and email on [Autopilot](/glossary/autopilot) behind a per-post review gate — so a Spanish cut lands on the accounts and cadence that reach a Spanish-speaking audience while the English one runs its own track. Descript makes the video speak every language; Kompozy makes each version show up, in the right format, on every feed.

Quick takeaways

  • Descript dubs video into 30+ languages with AI voices, and now regenerates the speaker's mouth with generative AI lip sync so the face matches the new language.
  • The lip sync is generated, not rotoscoped — it encodes the face into a latent space, generates new mouth movements from the translated audio, and blends them over the untouched background.
  • It can translate on-screen text layers, not just the audio, so graphics and captions localize with the voiceover.
  • Dubbing and lip sync run on the paid Creator plan and above and consume AI Credits; Business and Enterprise add transcript-level correction and a Do Not Translate list.
  • Descript localizes one recording into many language versions; use Kompozy to clip, caption, reframe, and publish each version across the eight social platforms plus blog and email.

Frequently asked questions

How many languages can Descript dub a video into?

Descript can translate and dub a recording into more than 30 languages using native-sounding AI voices generated from the transcript, and it can now apply AI lip sync so the speaker's mouth matches the dubbed language. Confirm the current language list on descript.com, since it changes.

How does Descript's lip sync work?

Rather than rotoscoping, Descript uses generative AI: it encodes the original video and reference frames of the speaker's face into a learned latent space, generates new mouth movements driven by the translated speech, and blends them over the untouched background and upper face — preserving the speaker's identity, lighting, and teeth.

What plan do I need for Descript dubbing and lip sync?

Dubbing and lip sync are available on Descript's paid Creator plan and above and consume AI Credits. On the Creator plan translation is automatic; Business and Enterprise plans let you fine-tune the translation in the transcript before it renders and use a Do Not Translate list to keep brand terms unchanged. Check descript.com/pricing for current details.

How do I distribute a dubbed video across platforms?

Descript localizes one recording into multiple language versions but publishes one project at a time. To turn each dubbed master into per-platform, per-language posts — captioned vertical clips, carousels, a blog, and a newsletter — and schedule them across the eight social platforms plus blog and email, creators pair a Descript master with a generation-and-publishing engine like Kompozy.

Related news

← All AI news · Get started →