// AI NEWS · FEATURE

Synthesia's AI Video Translator Dubs External Footage With Voice Cloning and Lip Sync

Synthesia's translation workflow dubs any uploaded clip or public YouTube link — cloning each speaker's voice and re-syncing lips into 140+ languages, no reshoot and no rebuild inside Synthesia required.

2026-09-23 · by Moe Ameen

What happened

Synthesia, the London-based AI avatar-video company, offers a video translator that works on footage that was never made inside Synthesia. You upload an MP4, MOV, or WEBM file — or paste a public YouTube URL — and the tool transcribes the audio, clones each speaker's voice, translates the script, and re-renders the video dubbed into the target language with the lips re-synced to the new audio. That sets it apart from tools whose lip-synced dubbing only works on video generated by their own avatars.

The workflow handles the parts that usually make dubbing painful. It detects and clones multiple speakers automatically with no manual voice assignment, keeps the dub matched to the correct on-screen speaker across cuts, and offers two lip-sync modes — Adaptive, which nudges speech speed to fit the new language, and Original, which preserves the source playback speed. Subtitles are auto-generated for every dubbed version and viewers can toggle them in Synthesia's Multilingual Player. Synthesia says the translator covers 140+ languages and regional variants, and can generate every language version in one pass rather than a separate job per language.

Voice handling is the other headline. The clone captures tone, emotion, pacing, and speaking style from the source audio and carries it into the target language, so a translated speaker sounds like themselves rather than a generic narrator. Creators who prefer not to clone can swap in a stock voice from Synthesia's library instead. A free tier lets you translate a first minute of video with full functionality; paid plans lift the limits and remove watermarks.

That's a meaningful distinction in the category: most AI video tools only localize material created inside their own editor — anything shot on a camera, cut in Premiere, or exported from a webinar platform typically has to be rebuilt before it can be translated. Dubbing external footage directly removes that rebuild step.

Why it matters for creators

  • A back catalog of existing videos — webinars, YouTube uploads, course modules, ads — becomes localizable without re-editing or recreating anything inside a specific tool.
  • Voice cloning that preserves each speaker's tone and pacing means a founder or host stays recognizable in every language, which matters for trust and brand continuity.
  • Automatic multi-speaker detection makes panel discussions, interviews, and podcasts dubbable — formats that manual voice-assignment workflows made tedious.
  • A free first minute lowers the barrier to testing whether AI dubbing is good enough for your audience before committing to a plan.
  • The bottleneck shifts from producing translations to distributing them: once you have a clip in 20 languages, you still need to cut, caption, and publish each one per market.

How to act on this with Kompozy

The translator solves localization; it does not solve distribution, and that is where the work actually piles up. Say you drop a 40-minute webinar recording into Synthesia and get it back dubbed into ten languages. You now hold ten long files and a blank calendar. Kompozy is built to close that gap. Bring each dubbed export in as source and Kompozy's Clipped Shorts turns the long recording into vertical cuts, auto-captions them in the dubbed language, reframes each for TikTok, Reels, and Shorts, and — from the same source — spins out Text Posts, Carousel Posts, a Blog Article, and an Email Newsletter so one localized asset becomes a full content week rather than a single upload.

Then autopilot publishes the set across the eight social platforms plus blog and email, on a schedule per market, through the per-post review pipeline so you approve before anything ships. The practical move today: localize your best-performing evergreen videos in Synthesia's translator, then run each language version through Kompozy to generate and publish a market-specific stream instead of manually re-cutting ten videos by hand.

Quick takeaways

  • Synthesia's video translator dubs uploaded files and public YouTube links, not just Synthesia-made avatar videos.
  • 140+ languages, multi-speaker voice cloning, lip re-sync, and auto subtitles, with a free first minute.
  • Localization is now cheap; per-market distribution is the remaining bottleneck — the part a content engine like Kompozy handles.

Frequently asked questions

Does Synthesia's video translator work on videos I didn't make in Synthesia?

Yes. You can upload an MP4, MOV, or WEBM file or paste a public YouTube URL. Synthesia transcribes the audio, translates it, clones the speakers' voices, and re-syncs the lips — the video does not need to have been created inside Synthesia.

How many languages does it support?

Synthesia says the translator covers 140+ languages and regional variants, and can generate all requested language versions in a single pass rather than running a separate job per language.

Does it clone the original speaker's voice?

Yes. The voice clone captures tone, emotion, pacing, and speaking style from the source audio and applies it in the target language, and it detects and clones multiple speakers automatically. You can also substitute a stock voice if you prefer.

Is it free?

There is a free tier that translates a first minute of video with full functionality. Paid plans raise the usage limits, add more languages and features, and remove the watermark.

Related news

← All AI news · Get started →