HeyGen's image-to-video avatar model — turn a single photo and a script into a talking video with hand gestures and voice-synced emotion.
Last verified · 2026-08-14 · by Moe Ameen
HeyGen Avatar IV is an image-to-video avatar model: you give it a single photo and an audio track or script, and it generates a video of that person talking and moving — synced lips, expressive facial movement, and, distinctively, authentic hand gestures. It is the model behind much of what people mean when they talk about "a talking photo": no camera, no motion capture, no video of the subject required, just one still image and words to say.
HeyGen introduced Avatar IV in 2025 as its most advanced image-to-video model at the time, and it pushed past earlier avatar generations in a few concrete ways. It can work from tilted heads, profiles, and angled source photos rather than demanding a clean front-facing shot; it reads tone and emotional content in the script to drive voice-synced expression; and it supports a range of styles from hyper-realistic human clones to stylized characters, including anime and animal/pet avatars, in portrait and fuller-body framings. HeyGen later exposed it through an Avatar IV API so developers can generate these talking videos programmatically from a photo and a script.
Under the hood it is a large model. In an August 13, 2026 engineering write-up with Google Cloud, HeyGen described Avatar IV as running on more than 18 billion parameters across three parts — a diffusion transformer for motion, a super-resolution transformer, and a VAE decoder for the final pixels — outputting 720p or 1080p at 25 frames per second. That post detailed porting the model to Google Cloud Trillium TPUs for a 1.86x speedup and up to 25% lower cost per generated minute. Note that HeyGen also ships newer and older avatar tiers — Avatar V is its newest, most realistic model, while the older Avatar III is cheaper per minute — so Avatar IV sits in the high-realism band of a lineup. Credit costs and model availability shift, so treat any exact figure as a snapshot of the official site.
The unlock in Avatar IV is the input: one photo becomes a presenter. That is a superpower for a specific problem — building a face-consistent, recurring brand personality — and it is exactly where [Kompozy](/) picks up. Avatar IV hands you a single talking clip; Kompozy turns that same identity into a persona that shows up, on brand, across every format and every feed. Its AI Influencer persona pool holds any number of personas with one primary identity, and Kompozy runs Avatar-IV-class generation natively inside [Persona Shorts](/glossary/persona-shorts) (avatar take plus auto-captions plus optional B-roll) and its longer-form Persona HeyGen format. Then [Persona Frames](/glossary/persona-frames) composites that avatar as a movable layer inside a pixel-exact [HyperFrames](/glossary/hyperframes) brand template, so the talking photo lands inside your actual design system instead of on a plain background.
From there the one photo stops being one clip. Governed by a [Persona Brief](/glossary/persona-brief) that fixes voice and banned phrases, the same persona fans out into carousels, quote cards, Persona Tweets, a blog article, and an email newsletter, and Kompozy auto-captions and reframes the video to 9:16, 1:1, and 16:9 before scheduling and publishing across Instagram, TikTok, YouTube, LinkedIn, X, Facebook, Pinterest, and Threads plus blog and email. Avatar IV makes a photo talk once; Kompozy makes that person your recurring, multi-platform presence.
Avatar IV is HeyGen's image-to-video avatar model. From a single photo plus an audio track or script, it generates a video of that person talking and moving, with synced lips, expressive facial movement, and authentic hand gestures. It launched in 2025 and can work from tilted, profile, or angled photos in realistic or stylized styles.
They are HeyGen avatar models at different realism and cost tiers. Avatar V is the newest and most realistic; Avatar IV is a high-realism image-to-video model that added authentic hand gestures and voice-synced emotion; Avatar III is older and cheaper per minute. The newer, more realistic models generally consume more credits per minute of video.
A single image of the person or character and something for them to say — an audio track or a text script. Avatar IV does not require video of the subject or motion capture. It can even generate from an angled or profile photo rather than a clean front-facing shot.
No. Avatar IV generates the talking video inside HeyGen; it does not caption it for silent autoplay, reframe it per feed, or schedule it. Kompozy runs Avatar-IV-class avatar generation in its Persona Shorts and Persona HeyGen formats, then auto-captions, reframes, repurposes, and publishes across nine platforms from one queue.
That is a natural fit for Kompozy. Its AI Influencer persona pool and Persona Brief keep one face and voice consistent, and it generates Avatar-IV-class video plus carousels, tweets, blogs, and newsletters from that identity — so a single photo becomes a persona posting on brand across every platform, not just one clip.