An honest 2026 review of Tavus Sparrow-2, the real-time turn-taking model: what it does well, its real limits, and who it actually fits.
Sparrow-2 is a genuinely strong real-time turn-taking model — Tavus reports #1 TurnBench dev-split results and a ~2.1% conversation failure rate, and in practice it makes live conversational avatars feel far less interruptive. Two honest caveats: it is proprietary and coupled to the Tavus stack, and it is not a content tool or a noise-canceller. Rate it as infrastructure for live voice/video agents, nothing more.
Sparrow-2 is Tavus's real-time conversational-understanding model, announced August 27, 2026 and live inside Tavus's stack. Its one job is turn-taking: deciding, moment to moment, whether the person talking to an AI has finished, is just pausing to think, or is interrupting. That sounds narrow, and it is — but it is also the single thing most voice AI still gets wrong, which is why it is worth reviewing on its own terms.
First, worth being precise about the category: Sparrow-2 is not an open-source noise-cancellation tool. It is proprietary, and it does not remove noise. It reasons over the whole audio scene — tone, pauses, speaker identity, backchannels, interruptions, and ambient sound together — roughly every 10 milliseconds, and Tavus calls this a new approach to the "cocktail party problem": understand the room instead of scrubbing it. If you arrived expecting a denoiser or a downloadable model, this is neither.
This review scores Sparrow-2 as what it is: infrastructure for real-time conversational agents. It is not a repurposing engine, a video generator, or a publishing tool, so we do not penalize it for lacking those — but we are explicit about the category line, because a lot of people find these pages hoping to make publishable content and Sparrow-2 will not do that for you.
Sparrow-2 is the turn-taking / conversational-understanding model in Tavus's family (which also includes the Phoenix video-rendering and Raven emotional-intelligence models). It sits inside Tavus's Conversational Video Interface (CVI) and PAL agents, and it is reachable through the Tavus API and the no-code PAL Maker. Rather than the traditional pipeline of voice-activity detection plus a fixed silence timeout, it treats the incoming audio stream as continuous evidence and predicts the conversational state at a native 10ms frame rate — faster than Sparrow-1's 40ms — so an agent can react to a barge-in or hold through a thinking pause more naturally. Tavus positions it as state-of-the-art on TurnBench, an independent, triple-annotated benchmark with a public leaderboard. On the dev split Tavus reports the field's highest end-of-turn recall (about 92%) and interruption recall (about 97%), and a conversation failure rate near 2.1% that it describes as roughly 4x lower than the best alternative it tested. Those headline figures are Tavus-reported; the benchmark itself is third-party, but treat the specific numbers as vendor claims until independently reproduced.
Sparrow-2 is for teams building real-time voice or video agents — conversational avatars, AI receptionists and phone agents, interactive demos, sales-page front-ends — who are already on or willing to adopt the Tavus stack, and for whom natural turn-taking in noisy, multi-speaker settings is a core requirement. It is not for creators who want to produce and publish content: it generates nothing you can post. And it is a poor fit for anyone who needs an open, portable, or self-hostable turn detector, or an actual noise-cancellation step in an audio pipeline.
| Dimension | Score | Why |
|---|---|---|
| Turn-taking accuracy | 4.7 / 5 | Tavus reports top-of-field end-of-turn and interruption recall on TurnBench; in live use it holds through thinking pauses and catches barge-ins well. |
| Real-time latency | 4.6 / 5 | Native 10ms frame rate (up from Sparrow-1's 40ms) gives fast, fine-grained reactions suited to live conversation. |
| Noise & multi-speaker robustness | 4.4 / 5 | Reasons over background speech and ambient noise rather than cancelling it — its stated strength on the "cocktail party" problem. |
| Benchmark transparency | 4.0 / 5 | TurnBench is independent with a public leaderboard, but the headline recall and failure-rate figures are Tavus-reported on the dev split. |
| Developer integration | 4.2 / 5 | Clean access via the Tavus API, PAL Maker, and CVI, with documentation — but only inside the Tavus ecosystem. |
| Openness & portability | 2.5 / 5 | Proprietary and coupled to the Tavus stack; there is no open-source model or standalone self-host option. |
| Pricing clarity | 3.5 / 5 | Sold as part of Tavus usage rather than a clearly published standalone price; confirm current rates on Tavus's pricing page. |
| Content-creation usefulness | 1.5 / 5 | Not its job — it renders no media and publishes nothing; scored low only to flag the category mismatch, not as a defect. |
Sparrow-2 is not sold as a standalone product with a public per-unit price. It ships as the turn-taking layer inside Tavus's Conversational Video Interface and PAL agents, so its cost is folded into Tavus's usage-based conversational-video pricing rather than a line item you buy on its own. That makes a clean cost comparison hard — you are pricing the whole real-time agent stack, not a turn detector.
For teams already committed to Tavus, that bundling is reasonable: you get turn-taking, perception, and avatar rendering as one integrated system, which removes the glue work of stitching a separate turn model into your pipeline. For teams that want just a turn detector to drop into an existing stack, the lack of a standalone, portable option is the real cost — you are buying into Tavus, not licensing a component.
Because the pricing is usage-based and can move, do not rely on any specific figure quoted secondhand. Confirm current rates and included conversation minutes on Tavus's own pricing page before budgeting, and model it as part of your total agent cost, not in isolation.
| Use case | Fit | Why |
|---|---|---|
| Real-time conversational avatars and AI agents | Strong | This is exactly what Sparrow-2 is built for — natural turn-taking is what makes a live agent feel human. |
| AI phone/voice receptionists and support | Strong | Correctly waiting through pauses and handling interruptions is the core UX problem here. |
| Noisy or multi-speaker environments | Strong | Reasoning over the whole scene is its stated advantage over VAD-plus-timeout pipelines. |
| A portable, open, or self-hosted turn detector | Weak | It is proprietary and tied to the Tavus stack — no open model or standalone deployment. |
| Standalone noise cancellation / audio cleanup | Weak | It understands noise to time turns; it does not remove noise from a recording. |
| Producing avatar/short-form content to publish | Weak | Wrong category — it generates no media and publishes nothing; use a content engine like Kompozy. |
| Podcast or recording cleanup in post | Weak | It is a real-time live-conversation model, not an offline audio editor. |
Kompozy is not a competitor to Sparrow-2 — it is on the other side of the line, and it is worth being clear about that so you pick the right tool. Sparrow-2 governs a live, two-way conversation: an avatar that listens and knows when to respond. Kompozy is a content generation and multi-platform publishing engine — it produces scripted, on-brand content you distribute one-to-many. It does not do real-time turn-taking, and Sparrow-2 does not produce or publish anything.
Where they meet is AI avatars. If you build a Sparrow-2-powered agent and record a demo or a live call, Kompozy clips and captions that footage into short-form for TikTok, Reels, and Shorts. And for net-new marketing you never have to film, Kompozy's Persona Shorts generate a talking-head avatar reading your script, then schedule it across the eight social platforms plus blog and email. Choose Sparrow-2 to make an agent converse; choose Kompozy to make and publish the content around it.
If you are building a real-time voice or video agent and are on (or open to) the Tavus stack, yes — turn-taking is the thing most voice AI gets wrong, and Sparrow-2 is reportedly best-in-field at it on the independent TurnBench benchmark. If you want an open, portable turn detector, a noise-canceller, or a way to make content, it is the wrong tool.
No. Sparrow-2 is proprietary. It is available only through Tavus — the Tavus API, the no-code PAL Maker, and the Conversational Video Interface — with no open-source release or standalone self-host option.
No. It reasons over background noise rather than removing it. Sparrow-2 models ambient sound, tone, pauses, and speaker identity together to decide when an AI should listen, wait, or speak. It is not a noise-cancellation or audio-cleanup tool.
Tavus reports it ranks #1 on TurnBench's dev split, with roughly 92% end-of-turn recall and 97% interruption recall, and a conversation failure rate near 2.1% that it calls about 4x lower than the best alternative it tested. TurnBench is independent with a public leaderboard, but treat the specific figures as vendor-reported until reproduced.
VAD plus a fixed silence timeout just detects whether someone is speaking and guesses they are done after a set pause. Sparrow-2 predicts the actual conversational state — done, thinking, or interrupting — by reading the whole scene every 10ms, which is why it handles thinking pauses and interruptions far better than a timeout.
No. It is infrastructure for live conversational agents and generates no media. To produce and publish avatar video, clips, carousels, blogs, or newsletters, pair it with a content engine like Kompozy, which generates those formats and distributes them across nine destinations — the eight social platforms plus blog and email.
For turn-taking specifically, open options like LiveKit's turn-detection model or Pipecat Smart Turn trade Tavus's integration for portability, and plain VAD-plus-timeout is the cheapest baseline. For the different job of producing and publishing content around your agent, Kompozy is the engine.
See Sparrow-2 (Tavus) vs Kompozy comparison → · Get Started →