Real-time, talk-back AI avatars you embed in a product or site — a photorealistic digital twin that listens and answers live, as the September 2026 TechCrunch journalist avatar demonstrated.
Last verified · 2026-09-26 · by Moe Ameen
Synthesia Interactive Avatars are the conversational, two-way side of the Synthesia platform: instead of rendering a fixed talking-head video from a script, they produce a photorealistic avatar that joins a live session, listens to a spoken question, and answers in real time. They jumped into public view on September 26, 2026, when TechCrunch reporter Dominic-Madori Davis published a first-person account of building an avatar of herself — from a photo session and roughly a two-minute voice recording, Synthesia produced four versions, including an interactive one you can actually talk to.
Technically, an interactive avatar is a chain of models — voice-to-text, a language model, and text-to-voice — feeding Synthesia's proprietary real-time video model. It's built for developers: you attach the avatar to your own agent, bring your own LLM and conversation logic, and can swap in third-party providers such as Cartesia, ElevenLabs, Google, or OpenAI for parts of the stack. Synthesia's docs describe the avatar joining a live video room (via LiveKit), and the library exceeds 240 avatars with broad language support, plus custom avatars of a specific person like the journalist twin. Pricing is pay-as-you-go at $0.12 per minute of conversation.
The single most instructive detail is grounding. Davis's interactive avatar was trained on one article — her reporting about venture-backed startup fraud — and constrained to answer only questions about that story, steering everything else back to the source. That scoping is what keeps a conversational likeness accurate instead of improvising. It also surfaces the trust question the piece is really about: she had to consent to the avatars being made, and audiences need to know when they're talking to an AI rather than a person.
The honest framing for a creator: this is a live agent you embed, not content you post. Interactive Avatars don't produce clips, carousels, images, blogs, or newsletters, they don't size anything for a feed, and there's no scheduler or publishing. They answer questions in real time — a different job from filling a content calendar.
Interactive Avatars are the piece Synthesia built for a live conversation — a concierge on your site that answers a visitor by voice. That's genuinely useful, and it's also the opposite half of a content operation: an interactive avatar waits for someone to arrive and ask; it never goes out and earns the audience. There's nothing to export, clip, or schedule, because live conversation isn't a file. So the useful pairing isn't "repurpose the avatar" — it's "run the outbound engine that fills the feed and points people at it." That outbound engine is [Kompozy](/).
Think of it as inbound and outbound sharing one identity. On the inbound side, Synthesia's interactive twin fields live questions. On the outbound side, Kompozy generates the recurring, published presence that gets anyone to show up in the first place: [Persona Shorts](/glossary/persona-shorts) and HeyGen Video Agent avatar video from a face-locked persona, plus [Clipped Shorts](/glossary/persona-shorts), [Carousel Posts](/glossary/hyperframes), Photo Posts, quote graphics, Blog Articles, and Email Newsletters — every asset held to one [Persona Brief](/glossary/persona-brief) so the face and voice on your feed match the twin answering on your site. Then [Autopilot](/glossary/autopilot) schedules and publishes the set across the eight social platforms plus blog and email behind a per-post review gate, which is also where you add a clear AI-disclosure — the exact safeguard the journalist experiment says matters most. Synthesia's avatar handles the conversation when someone lands; Kompozy is what makes them land, on brand, every week.
It's a photorealistic AI avatar that holds a real-time voice conversation — it listens to a spoken question and answers on the fly, rather than playing a pre-rendered video. It runs on a voice-to-text, LLM, and text-to-voice stack over Synthesia's real-time video model, and can be grounded to a specific knowledge base. The September 2026 TechCrunch journalist avatar, trained on one article, is a public example.
Not directly — a live conversation isn't an exportable file, and Interactive Avatars don't generate clips, carousels, images, or a video library, and don't publish anything. To build a published avatar presence, you use a generation-and-distribution engine: Kompozy creates avatar video plus other formats from a face-locked persona and schedules them across nine platforms, while the Synthesia twin handles live Q&A.
Synthesia lists Interactive Avatars on pay-as-you-go pricing at $0.12 per minute of conversation, with no minimum commitment. Custom-avatar creation and enterprise terms sit alongside Synthesia's broader plans and change often, so confirm current figures on Synthesia's site before building on it.
Ground it. Synthesia's journalist demo trained the avatar on a single article and had it refuse off-topic questions — a scoped knowledge base is what stops a conversational likeness from improvising answers it can't stand behind. Pair that with clear disclosure that the avatar is AI, which the reporter's own conclusion frames as the load-bearing trust safeguard.
No — Kompozy has no real-time conversational avatar; that's Synthesia's job. Kompozy does the opposite half: it generates and publishes on-brand avatar video and other content from a persona you control across nine platforms. The two pair cleanly — Synthesia's twin for live conversation, Kompozy for the outbound content that builds the audience.