Descript's video-translation feature: dub a recording into 30+ languages with AI voices, then regenerate the speaker's mouth so the face matches the dubbed language — plus on-screen text-layer translation.
Last verified · 2026-08-17 · by Moe Ameen
Descript's translation and dubbing is the localization layer inside the transcript-based editor: it turns one recording into many language versions. Descript transcribes the file, translates that transcript, generates a native-sounding AI voiceover in the target language, and re-times it to the video. It supports more than 30 languages, and because the whole thing is driven by the transcript, you correct a translation by editing text rather than re-recording audio.
The addition that makes it stand out is AI lip sync. After the dubbed voiceover is generated, Descript automatically regenerates the lower half of the speaker's face so the mouth moves to the new language's sounds and pacing. It is generative, not rotoscoped: Descript encodes the original video and reference frames of the speaker's face into a learned latent space, generates new mouth movements from the translated speech, and blends them over the untouched background and upper face — preserving the speaker's identity, lighting, and teeth. It also translates text layers, so on-screen graphics and captions localize alongside the audio.
Dubbing and lip sync run on Descript's paid Creator plan and above and consume AI Credits, and longer clips take longer to render. On the Creator plan translation is automatic; Business and Enterprise plans add human-in-the-loop control — you can fine-tune the translation directly in the transcript before it renders, and apply a Brand Studio "Do Not Translate" list so product names and brand terms stay intact through localization. Descript has documented the workflow in its help center and described the engineering behind multilingual dubbing at scale in a case study with OpenAI. Confirm the current language list, plan gates, and credit costs on descript.com, since AI features change often.
The honest boundary: this feature is localization, not distribution. It produces watchable language versions of one recording. It does not clip those versions into per-platform posts, build carousels or a blog from them, or schedule and publish them across every network — that downstream step is a separate job.
Descript's dubbing gives you something rare: several believable language versions of the same talking-head video, each with the mouth actually moving to the words. But a dubbed MP4 is still a single asset — and the value of localizing is only realized if each version becomes a full content presence in that language. That is what [Kompozy](/) adds, and it is deliberately the opposite half of the problem: Kompozy does not dub or lip-sync (Descript owns that), and Descript does not manufacture the surrounding content set or publish it everywhere (Kompozy owns that).
Feed a Descript language master into Kompozy and it generates the formats a dubbing tool can't: [Clipped Shorts](/glossary/clipped-short) auto-cut from the long video with word-synced captions, brand-exact [Carousel Posts](/glossary/hyperframes) and [Quote Graphics](/glossary/output-buckets), an [Infographic Photo](/glossary/output-buckets), a [Blog Article](/glossary/output-buckets), and an [Email Newsletter](/glossary/output-buckets) — every output reframed to 9:16, 1:1, and 16:9 and held to one voice by the [Persona Brief](/glossary/persona-brief). Then it schedules and publishes across the eight social platforms plus blog and email with a per-post review pipeline and [Autopilot](/glossary/autopilot). So the pipeline is clean: Descript localizes one recording into many languages; Kompozy turns each localized cut into a week of platform-native posts and pushes them live. One is the translator, the other is the factory and the distributor.
It is Descript's translation-and-dubbing feature: it transcribes a recording, translates the transcript, and generates a native-sounding AI voiceover in the target language, supporting more than 30 languages. With AI lip sync on, it also regenerates the speaker's mouth so the face matches the dubbed language.
Descript uses generative AI, not rotoscoping. It encodes the original video and reference frames of the speaker's face into a learned latent space, generates new mouth movements driven by the translated speech, and blends them over the untouched background and upper face, keeping the speaker's identity, lighting, and teeth consistent.
Dubbing and lip sync are available on Descript's paid Creator plan and above and consume AI Credits. The Creator plan translates automatically; Business and Enterprise let you correct the translation in the transcript and use a Do Not Translate list. Check descript.com/pricing for current numbers.
Yes — beyond the voiceover, Descript can translate text layers, so on-screen graphics and captions are localized alongside the audio rather than left in the source language. Confirm the current behavior in Descript's help center.
Descript localizes one recording into multiple language versions with matched mouths; Kompozy turns each dubbed master into platform-native content — captioned clips, carousels, a blog, and a newsletter — and publishes it across the eight social platforms plus blog and email. Descript is the translator; Kompozy is the generation-and-distribution engine.