// HOW-TO · AI VIDEO

How to adjust HeyGen Voice settings for a natural, on-brand read (2026)

How to tune HeyGen Voice settings: pick the right voice and engine, set speed and pitch, shape expression, add pauses, match accent, and clone a voice with consent.

Last verified · 2026-10-10 · by Moe Ameen

HeyGen's voice layer has a lot more control than the single "pick a voice" dropdown most people stop at. Between the voice catalog, the per-script sliders in the app, the `voice_settings` object in the API, and the controls on HeyGen's own in-house HeyGen Voice model, you can set speaking speed, pitch, expression, pauses, and accent — and build a cloned voice that sounds like a specific person. Left at defaults, a HeyGen read can land flat or rushed; tuned, it carries the emphasis and pacing that make an avatar video feel like a person talking.

This walks through the settings in the order you actually touch them: choose the voice and check what it supports, set speed and pitch, shape the emotional range, add pauses and accent hints, then — if you want the read to sound like you or a consented speaker — create a voice likeness with HeyGen Voice. Controls differ by voice engine and by whether you're in the app or the API, so the first step is always confirming which knobs the voice you picked actually exposes. Note that HeyGen Voice (the in-house model) and its API endpoints were in preview as of late 2026, so some controls require your account to be enabled first — verify current availability in HeyGen's own docs before building around a specific parameter.

The steps

  1. Pick the voice — and check what it supports before you tune. HeyGen's catalog spans 300+ voices across several engines (its own HeyGen Voice, plus third-party engines). In the API, `GET /v3/voices` lists them with filters for language, gender, and engine; each voice returns flags like `support_pause` (whether `<break>` pause tags work) and `support_locale` (whether you can push an accent). Not every control works on every voice, so read those flags first — tuning a parameter a voice doesn't support silently does nothing.
  2. Set speaking speed and pitch. These are the two most reliable controls and they exist on nearly every voice. In the API, the `voice_settings` object on video generation takes `speed` (0.5 to 1.5, where 1.0 is normal) and `pitch` (-50 to +50 semitones); in the app, the same sit as per-script sliders. Nudge speed down slightly for authority or up for energy, and leave pitch alone unless the default read sits noticeably high or low — large pitch shifts are the fastest way to make a voice sound synthetic.
  3. Shape the expression — and change one slider at a time. How you control emotion depends on the voice. For HeyGen Voice Instant clones, the control is `expressiveness_boost`. For HeyGen Voice Professional clones, you get `pitch_variance` (how much the pitch moves across a sentence) alongside `speed`, `pitch_shift`, and a `seed`. For third-party (ElevenLabs-engine) voices, the app exposes Stability, Similarity, Style, and Speaker Boost sliders — lower Stability widens emotional range, higher flattens it. Whatever the engine, move one slider at a time and re-render; changing several at once makes it impossible to tell which one helped.
  4. Add pauses and emphasis with break tags. On any voice whose `support_pause` flag is true, insert `<break>` tags in the script to add deliberate silence — a short beat before a key line reads far more naturally than the engine's default run-on pacing. Keep pauses purposeful; a break after every clause makes the delivery stilted. If a `<break>` tag shows up spoken aloud or ignored in the output, the voice you picked doesn't support it — switch to one that does.
  5. Match the accent or language with a locale hint. For multilingual voices (those with `support_locale`), set a `locale` value (a BCP-47 code like `en-GB` or `es-MX`) to steer the accent and regional pronunciation without changing the voice itself. This is how you keep one persona's voice but localize the read per market. On voices that don't carry the locale flag, the hint is ignored, so confirm support before you rely on it for a translated video.
  6. Create a voice likeness — Instant or Professional, with consent. To make the voice sound like a specific person, HeyGen Voice offers two clone modes. Instant builds a voice from a single short recording and is usually ready within minutes but can't be retrained — you make a new one to change it. Professional trains on 1-10 recordings totaling at least ~20 minutes for a closer match, takes longer, and can be retrained. Either way, cloning a real person's voice requires that person's explicit consent — HeyGen gates likeness creation behind an explicit consent step for exactly this reason.
  7. Preview, then lock the settings to the voice you reuse. Render a short test line and listen on the device your audience uses — phone speakers expose harshness a laptop hides. Once a combination of speed, pitch, expression, and pauses reads the way you want, treat it as the standard for that voice rather than re-tuning per video. Consistency is what makes a recurring avatar feel like one presenter instead of a slightly different narrator each week.

Common gotchas

  • Controls vary by voice and engine. `expressiveness_boost` is a HeyGen Voice control; Stability/Similarity/Style belong to third-party ElevenLabs-engine voices. Check the voice's flags and engine before expecting a given slider to exist.
  • A `<break>` tag only works where `support_pause` is true. On an unsupported voice it's ignored or read aloud — confirm the flag before scripting around pauses.
  • Pushing Style high (on ElevenLabs-engine voices) makes the model less stable; the standard advice is to keep it near 0 and get range from Stability instead.
  • High Similarity on a noisy source recording reproduces the noise. Clone from clean, unprocessed audio or you bake hiss and room tone into every render.
  • Professional clones consume a purchased voice slot. Generating speech on a voice beyond your slot limit returns a `voice_expired` error until you free or add a slot.
  • Instant clones can't be retrained. If the Instant voice isn't right, you create a new one — you don't edit the existing clone.
  • Big speed or pitch moves backfire. Past a small nudge, they're the clearest tell that a voice is synthetic; fix a flat read with expression and pauses first.
Legal note

Cloning a voice that isn't your own — a colleague, a client, a public figure — requires that person's explicit, documented consent, and HeyGen requires an explicit consent step before it will build a likeness. Using a real person's voice without permission can violate right-of-publicity and anti-impersonation laws (and platform policies) regardless of how the clone was made. Separately, most platforms require AI-generated or AI-altered media to be labeled; a tuned, lifelike synthetic voice is exactly the kind of content those disclosure rules target, so apply the relevant label when you publish.

Where Kompozy fits

Every control above is a per-render chore: you tune speed, pitch, expression, and pauses, render, listen, and repeat — then do it again on the next video, hoping you matched last week's settings. That's fine for a one-off, but it's the wrong unit of work if you publish a recurring presenter. [Kompozy](/) changes the unit. You tune a HeyGen voice once, bind it to an [AI Influencer persona](/glossary/avatar-video), and every render after that speaks with the same dialed-in read — no slider-fiddling per clip, because the voice is a property of the persona, not a setting you re-enter each time.

That binding is why Kompozy treats HeyGen's voice as an ingredient rather than a destination. [Persona Shorts](/glossary/persona-shorts), [Persona Frames](/glossary/persona-frames), and the longer-form Persona HeyGen format all generate HeyGen avatar video using that persona's locked voice and face, and a [Persona Brief](/glossary/persona-brief) governs the written side — captions, blogs, newsletters — so the voice your audience hears and the voice they read stay the same identity. HeyGen's settings panel tunes one voice track; Kompozy carries that tuned read across [18 output formats](/glossary/output-buckets).

Then it ships. The thing HeyGen's voice controls can't do — reframe for muted feeds, auto-caption, and schedule — is the back half Kompozy owns: it fans the finished, voiced video across the eight social platforms plus blog and email on [Autopilot](/glossary/autopilot), behind a per-post review gate where you confirm the AI label this tutorial's legal note calls for. Honest boundary: Kompozy doesn't replace HeyGen's voice model or its clone-creation flow — you still build the consented likeness in HeyGen — it's the engine that reuses that voice at scale and turns each read into published posts. Pricing is credit-based: Starter ($199/mo, 5,500 credits) suits a solo creator running one voiced persona, Pro ($499/mo, 18,000 credits) fits higher-volume output and teams, and Enterprise is custom.

Frequently asked questions

How do I make a HeyGen voice sound less robotic?

Start with pacing, not pitch. Slow the speed slightly from the 1.0 default, add a `<break>` tag before your most important line, and — depending on the voice engine — raise expression (lower Stability on ElevenLabs-engine voices, or raise `expressiveness_boost` on HeyGen Voice Instant clones). Change one setting per render so you can hear what each does. Large pitch shifts usually make a voice sound more synthetic, not less, so leave pitch near default.

What does expressiveness_boost do in HeyGen Voice?

It's the emotion control for HeyGen Voice Instant clones — raising it widens the voice's expressive range so the read carries more emphasis and feeling instead of a flat, even delivery. It's specific to the Instant mode of HeyGen's own voice model; Professional clones expose `pitch_variance` and a `seed` instead, and third-party voices use the Stability/Style sliders.

Can I add pauses to a HeyGen voiceover?

Yes, on voices that support it. Insert `<break>` tags in the script to add deliberate silence, which reads far more naturally than the default run-on pacing. It only works where the voice's `support_pause` flag is true — if a break tag is ignored or spoken aloud, pick a voice that supports pauses.

What is the speed and pitch range in HeyGen?

Through the API's `voice_settings`, speed runs 0.5 to 1.5 (1.0 is normal) and pitch runs -50 to +50 semitones. The app exposes the same as per-script sliders. In practice, small adjustments read best — a slight speed change for energy or authority, and pitch left near default unless the voice sits noticeably high or low.

Instant vs Professional voice clone — which should I use?

Use Instant for speed and testing: one recording, ready within minutes, no retraining (you make a new one to change it). Use Professional when the match matters: 1-10 recordings totaling at least ~20 minutes, longer to train, retrainable, and it consumes a purchased voice slot. Both require the voice owner's explicit consent.

Do I need consent to clone a voice in HeyGen?

Yes, if the voice isn't your own. HeyGen gates likeness creation behind an explicit consent step, and cloning someone else's voice without documented permission can breach right-of-publicity and anti-impersonation laws as well as platform terms. Clone your own voice freely; clone anyone else's only with their recorded consent.

Related tutorials

← All how-to guides · Get Started