Sparrow-2 is Tavus's new conversational-understanding model — it reads a whole audio scene every 10ms to decide when to listen, wait, or speak, and it reasons over background noise instead of cancelling it. It is not an open-source noise-cancellation tool.
2026-09-09 · by Moe Ameen
On August 27, 2026, Tavus — the conversational-video AI company behind the Phoenix, Raven, and Sparrow model family — announced Sparrow-2, a real-time conversational-understanding model built to handle turn-taking in live voice and video conversations. It is live now as a general-availability model inside Tavus's stack, reachable through the Tavus API, the no-code PAL Maker, and the Conversational Video Interface (CVI) that powers Tavus's real-time avatars. There is an interactive demo at sparrow2.tavuslabs.org.
Worth clearing up front: Sparrow-2 is not an open-source noise-cancellation tool. It is a proprietary turn-taking model, and its approach to noise is the opposite of cancellation. Instead of stripping background sound out of the signal, Sparrow-2 treats the entire incoming audio stream as evidence — modeling tone, pauses, speaker identity, backchannels ("mm-hm"), interruptions, and ambient noise together to decide, roughly every 10 milliseconds, whether the person is done talking, thinking, or being interrupted by the room. Tavus frames this as a new take on the "cocktail party problem": understand the whole scene rather than scrub it.
Tavus positions Sparrow-2 as state-of-the-art on TurnBench, an independent, triple-annotated benchmark of real human conversation with a public leaderboard and a field of roughly fourteen systems. On the benchmark's dev split, Tavus reports the highest end-of-turn recall in the published field (about 92%, well ahead of the runner-up) and the highest interruption recall (about 97%), and cites a conversation failure rate near 2.1% — which it describes as roughly 4x lower than the best alternative it tested. The model runs at a native 10ms frame rate (up from Sparrow-1's 40ms) for faster, finer reactions. Treat the exact figures as Tavus-reported until independently reproduced, but the direction — sharper, faster turn detection in noisy, multi-speaker settings — is the story.
There are two honest ways to act on this, and neither is "use Sparrow-2 to make posts" — it does not make content. The first is coverage. "Sparrow-2" and "Tavus turn-taking" are high-intent searches this week while builders work out what the model actually is and how it differs from noise cancellation. A single clear point of view turns into a week of content inside [Kompozy](/): a short explainer, a [Carousel](/glossary/hyperframes) breaking down turn-taking vs. noise cancellation, captioned [Clipped Shorts](/glossary/clipped-short) from a screen-recorded demo, a Blog Article, and platform-native [Text Posts](/glossary/output-buckets) — all held to one voice by the [Persona Brief](/glossary/persona-brief) and pushed across eight social platforms plus blog and email on [autopilot](/glossary/autopilot) while interest is still climbing.
The second is the workflow distinction the launch draws so cleanly. Tavus and Sparrow-2 own the live, two-way conversation — an avatar that listens and responds in real time. Kompozy owns the other half: producing scripted, on-brand avatar and short-form content you publish one-to-many. If you demo a Sparrow-2-powered agent on a call or a stream, Kompozy clips and captions that recording into shorts; and for net-new marketing you never have to film, [Persona Shorts](/glossary/persona-shorts) generate a talking-head avatar reading your script. Different problems, complementary tools — see the honest breakdown in our [AI video generators compared](/ai-content/video-generator-comparison), which already covers where Tavus wins.
No — that framing is inaccurate. Sparrow-2 is a proprietary real-time turn-taking (conversational-understanding) model from Tavus, available through the Tavus API, the no-code PAL Maker, and the CVI platform. It does not cancel noise; it reasons over background noise, tone, pauses, and speaker identity together to decide when an AI should listen, wait, or speak.
It handles turn-taking in live voice and video conversations. Roughly every 10 milliseconds it reads the whole audio scene — semantics, prosody, pauses, backchannels, interruptions, and ambient noise — to predict whether the person has finished a turn, is just pausing to think, or is interrupting. Tavus uses it to make its real-time conversational avatars feel less like they are talking over you.
No. It is infrastructure for real-time conversational agents, not a content generator. It writes no scripts, renders no video, and publishes nowhere. To produce and distribute avatar/short-form/written content, pair it with a content engine like Kompozy, which generates persona video, clips, carousels, blogs, and newsletters and publishes across the eight social platforms plus blog and email.