Descript's translation-and-dubbing now goes beyond swapping the audio: it uses generative AI to regenerate the lower half of the speaker's face so the mouth moves to the dubbed language, and it can translate on-screen text layers too — putting believable localized video within reach of a solo creator.
2026-08-17 · by Moe Ameen
Descript, the transcript-based AI audio and video editor, has expanded its translation and dubbing so a localized video can look as native as it sounds. The tool already dubbed a recording into more than 30 languages using native-sounding AI voices generated from the transcript. The expansion adds AI lip sync: after it produces the translated voiceover, Descript automatically regenerates the lower half of the speaker's face to match the new language's sounds, pacing, and mouth movements, so the video no longer reads as a dub laid over a mismatched mouth.
The lip sync is generative rather than rotoscoped. Descript encodes the original video and reference frames of the speaker's face into a learned latent space, generates new mouth movements driven by the translated speech, and blends them back over the untouched background and upper face — keeping the speaker's identity, lighting, and teeth consistent. It also translates text layers, so on-screen graphics and captions can be localized alongside the audio instead of being left in the source language.
Availability follows Descript's plan tiers. Dubbing and lip sync run on the paid Creator plan and above, and the feature consumes AI Credits; longer videos take longer to render. On the Creator plan the translation is automatic, while Business and Enterprise plans let you fine-tune the translation directly in the transcript before it renders and apply a "Do Not Translate" list from Brand Studio so product names and brand terms survive localization unchanged. Descript has documented the workflow in its help center and detailed the engineering behind multilingual dubbing at scale in a case study with OpenAI.
The framing to keep straight: this is localization of an existing recording — one polished video turned into many language versions with matched mouths. It is not, on its own, a system for cutting those versions into per-platform posts or distributing them across every network in each region. Confirm the current language list, plan gates, and credit costs on descript.com, since AI features and their limits change often.
Descript's lip-synced dubbing solves localization: one recording becomes a stack of language masters with matched mouths. It does not solve what comes after — turning each of those masters into per-platform, per-region posts and getting them out everywhere. That is the gap [Kompozy](/) is built for. Take a Descript master — the source cut and each dubbed version — into Kompozy, and it fans every one into the feed-native set: vertical [Clipped Shorts](/glossary/clipped-short) with word-synced captions in the matching language, a [Carousel](/glossary/hyperframes) and [Quote Graphics](/glossary/output-buckets) of the key points, a [Blog Article](/glossary/output-buckets) and an [Email Newsletter](/glossary/output-buckets) written from the transcript, each held to one voice by the [Persona Brief](/glossary/persona-brief).
The move a creator makes this week: localize the video once in Descript, then let Kompozy handle the multiplication. It reframes each language cut to 9:16, 1:1, and 16:9, generates the surrounding formats per version, and schedules and publishes across the eight social platforms plus blog and email on [Autopilot](/glossary/autopilot) behind a per-post review gate — so a Spanish cut lands on the accounts and cadence that reach a Spanish-speaking audience while the English one runs its own track. Descript makes the video speak every language; Kompozy makes each version show up, in the right format, on every feed.
Descript can translate and dub a recording into more than 30 languages using native-sounding AI voices generated from the transcript, and it can now apply AI lip sync so the speaker's mouth matches the dubbed language. Confirm the current language list on descript.com, since it changes.
Rather than rotoscoping, Descript uses generative AI: it encodes the original video and reference frames of the speaker's face into a learned latent space, generates new mouth movements driven by the translated speech, and blends them over the untouched background and upper face — preserving the speaker's identity, lighting, and teeth.
Dubbing and lip sync are available on Descript's paid Creator plan and above and consume AI Credits. On the Creator plan translation is automatic; Business and Enterprise plans let you fine-tune the translation in the transcript before it renders and use a Do Not Translate list to keep brand terms unchanged. Check descript.com/pricing for current details.
Descript localizes one recording into multiple language versions but publishes one project at a time. To turn each dubbed master into per-platform, per-language posts — captioned vertical clips, carousels, a blog, and a newsletter — and schedule them across the eight social platforms plus blog and email, creators pair a Descript master with a generation-and-publishing engine like Kompozy.