HeyGen moved past licensing third-party audio and shipped its own voice model on October 9, 2026 — free inside the platform and API, with a $99/month Professional Voice Clone add-on — pairing one identity's face and voice in the same Avatar V stack.
2026-10-10 · by Moe Ameen
On October 9, 2026, HeyGen announced HeyGen Voice, its first in-house voice model, and said it debuted at #1 on Artificial Analysis' independent voice leaderboard. Until now HeyGen's avatar videos leaned on third-party audio; HeyGen Voice brings the voice layer inside the company's own stack, so a video's face and voice are generated from one system rather than stitched from separate vendors.
HeyGen frames the model around identity. Its Avatar V model is the core, and voice is described as a new format inside the same end-to-end stack — the pitch is that the output keeps a speaker's tone, pacing, and expression instead of flattening it into a generic synthetic read. CTO Rong Yan put the positioning as "you don't have to trade authenticity for quality," pointing to the leaderboard placement as independent proof.
On pricing, HeyGen Voice is free within the HeyGen platform and API. A separate paid add-on, HeyGen Professional Voice Clone, costs $99 per month and is trained — with the voice owner's explicit consent — on roughly 30 minutes to three hours of their own speech. HeyGen says it requires that explicit consent and builds safeguards into its products so a person stays in control of how their voice and likeness are used.
The launch lands on a company that has had a loud 2026: founded in 2020, HeyGen now reports more than 40 million users across 196 countries, more than 118 million videos created, adoption across 85% of the Fortune 100, and over $200 million in ARR after doubling revenue in eight months. One figure HeyGen cited to explain the voice push: a survey of more than 1,000 small business owners found 71.6% had recorded a business video and then decided not to post it. Treat the leaderboard claim as a snapshot — independent voice rankings move, and the headline position is drawn from a cloned-voice comparison rather than every voice task — and confirm current specifics against HeyGen's own materials.
HeyGen Voice solves the "does this sound like me" half of avatar video, and that is real — a consistent voice is what makes a recurring persona believable. But a voice model outputs audio; it does not caption that audio for the 85% of feeds that autoplay muted, size it for nine destinations, or decide what gets posted when. That is the layer [Kompozy](/) owns. Its [Persona Shorts](/glossary/persona-shorts) format already renders a HeyGen talking-head avatar — voice and face together — and auto-captions it for silent scrolling, and [Persona HeyGen](/ai-tools/heygen) drives longer multi-scene cuts from your own persona pool, so the voice layer HeyGen just upgraded flows straight into finished, captioned video.
The difference shows up after the render. Kompozy fans one idea into a short, a carousel, a quote card, an X thread, a blog, and a newsletter, then schedules the whole set across the eight social platforms plus blog and email — so an on-brand voice becomes a full feed, not a single audio clip. [Autopilot](/glossary/autopilot) turns your source material into a recurring cadence behind a per-post review gate, and the [Persona Brief](/glossary/persona-brief) keeps the written copy in the same voice the audio speaks in. HeyGen gave the persona a better voice; Kompozy is what puts that persona in front of an audience on a schedule.
HeyGen Voice is HeyGen's first in-house AI voice model, announced October 9, 2026. It generates the voice layer inside HeyGen's own Avatar V stack — so a video's face and voice come from one system — and is designed to keep a speaker's tone, pacing, and expression rather than sound generically synthetic.
The base HeyGen Voice model is free within the HeyGen platform and API. A separate add-on, HeyGen Professional Voice Clone, costs $99 per month and is trained — with the voice owner's explicit consent — on roughly 30 minutes to three hours of their own speech.
HeyGen says it debuted at #1 on Artificial Analysis' independent voice leaderboard. That headline result is drawn from a cloned-voice comparison and reflects a point-in-time snapshot; independent rankings move, so the practical test is how the voice performs on your own script and in your language.
HeyGen generates the voice and avatar video; it does not caption it for muted autoplay, size it per platform, or schedule it. Kompozy does that — its Persona Shorts and Persona HeyGen formats render HeyGen avatar video, auto-caption it, repurpose it into other formats, and publish across eight social platforms plus blog and email behind a per-post review gate.