// AI NEWS · MODEL RELEASE

Nari Labs' Qwen3-TTS and Qwen3-ASR Top Coval's Voice-AI Benchmarks on Latency, Accuracy, and Cost

The team behind the open-source Dia voice model now leads Coval's public voice-AI benchmark on both text-to-speech and speech recognition — sub-70ms latency, low word-error rates, and pricing well under the incumbent voice APIs.

2026-09-14 · by Moe Ameen

What happened

On September 14, 2026, Nari Labs — the team behind Dia, an open-source dialogue text-to-speech model released under Apache 2.0 and covered as a challenger to incumbents like ElevenLabs — announced that its hosted Qwen3-TTS 1.7B and Qwen3-ASR 1.7B models rank at the top of Coval's public voice-AI benchmark. Coval runs a reproducible, open-source benchmark of speech-to-text and text-to-speech providers on a pinned dataset and publishes the results at benchmarks.coval.ai. Both Nari models are optimized servings of Alibaba's Qwen3 speech models, run on Nari's own low-latency inference stack.

On the speech-to-text side, Nari's Qwen3-ASR Fast placed first in latency and second in accuracy: a reported median time-to-final-segment around 44 ms and a word error rate of about 3.6%. On text-to-speech, Nari's Qwen3-TTS Fast placed second in latency and first in accuracy: a reported median time-to-first-audio around 63 ms and a word error rate of about 3.8%. Nari frames the result as sitting on the quality-latency Pareto frontier for both tasks — you don't trade accuracy for speed. For context, the benchmark recorded the official Qwen3 TTS Flash realtime endpoint at 8.8% WER and 692 ms median time-to-first-audio, so Nari's serving is both markedly more accurate and roughly an order of magnitude faster to first sound.

Price is the other half of the pitch. Nari lists its hosted TTS at $10 per 1M characters on the Fast tier and $5 per 1M on Standard, and its ASR at $0.12 per hour of audio on Fast and $0.06 on Standard — well below premium voice APIs (Nari cites ElevenLabs and Cartesia TTS pricing at roughly 5–6.5x higher per character). Nari had already open-sourced, in August 2026, an Apache-2.0 inference engine for Qwen3-TTS 1.7B that reached sub-50 ms p95 time-to-first-audio at 10 requests per second on a single NVIDIA H100, which works out to roughly $2 per 1M characters if you self-host at full utilization.

Treat the exact latency figures, word-error rates, and prices as a benchmark-day snapshot — voice benchmarks re-run on a schedule and provider pricing changes. Confirm current numbers on Coval's benchmark page and Nari's own pricing before you build on them.

Why it matters for creators

  • Real-time voice got cheap and fast at the same time. Sub-70 ms first-sound latency is built for live, interruptible voice agents; the pricing makes high-volume narration and transcription viable for solo creators, not just funded apps.
  • The headline gains are on the audio layer — synthesizing a voice, transcribing a recording. Turning that audio into captioned, reframed, on-brand video across every feed is a separate job the benchmark doesn't touch.
  • Cheap, accurate ASR is quietly the bigger creator unlock. A low word-error transcript of a podcast, livestream, or voice memo is the raw material a repurposing pipeline runs on — and at $0.06–$0.12 per hour, transcribing everything you record stops being a budget question.
  • Open-source lineage matters: Nari's serving stack for Qwen3-TTS is Apache-2.0, so the same speed is available to self-hosters, not locked behind one vendor's API.
  • Lower audio cost pushes the bottleneck downstream. When voice and transcription are near-free, the thing standing between an idea and an audience is production and distribution volume, not the price of a voiceover.

How to act on this with Kompozy

The useful read on this news isn't "switch your voice tool" — it's that the audio layer of content just got faster and markedly cheaper, while the part that actually builds an audience didn't move at all. A 3.8%-WER voiceover or a 3.6%-WER transcript is a file; it grows no following on its own. [Kompozy](/) is the engine that closes that gap, and the play here is to treat Nari's cheap, accurate audio as an input and let Kompozy do the making and the publishing around it.

Concretely, you can act on this today two ways. Feed a long recording — a podcast, a webinar, a rambling voice note — through cheap Qwen3-ASR to get a clean transcript, drop that transcript into Kompozy's Quick Ingest, and one source fans out under a [Persona Brief](/glossary/persona-brief) into [Clipped Shorts](/glossary/clipped-short), brand-exact carousels and quote graphics via [HyperFrames](/glossary/hyperframes), native text posts, a blog article, and an email newsletter. Or generate the written content in Kompozy first, narrate it with low-cost Qwen3-TTS, and pair the track with an avatar-narrated [Persona Short](/glossary/persona-shorts) or a faceless Listicle Video — Kompozy's own avatar formats already speak with HeyGen's native TTS, so you only reach for Nari when you want a standalone, dirt-cheap audio track. Either way, [Autopilot](/glossary/autopilot) schedules and publishes the whole set across the eight social platforms plus blog and email behind a per-post review gate. Nari made the sound cheap; Kompozy turns it into a week of posts. For the audio-into-video mechanics, see the [voice-cloning-for-video playbook](/how-to/voice-cloning-for-video-content).

Quick takeaways

  • Nari Labs' hosted Qwen3-TTS 1.7B and Qwen3-ASR 1.7B lead Coval's public voice-AI benchmark, announced September 14, 2026.
  • Qwen3-ASR Fast: ~44 ms median time-to-final-segment, ~3.6% WER — #1 latency, #2 accuracy. Qwen3-TTS Fast: ~63 ms time-to-first-audio, ~3.8% WER — #2 latency, #1 accuracy.
  • Pricing: TTS at $10 / 1M chars (Fast) and $5 / 1M (Standard); ASR at $0.12 / hr (Fast) and $0.06 / hr (Standard) — well under premium voice APIs.
  • Both are optimized servings of Alibaba's Qwen3 speech models on Nari's low-latency stack; Nari's Qwen3-TTS inference engine is Apache-2.0 and reached sub-50 ms p95 latency on a single H100.
  • The models make cheap, accurate audio and transcripts — not captioned, reframed, published content. Pairing them with a generation-and-publishing engine like Kompozy is what turns the audio into posts.

Frequently asked questions

What are Nari Qwen3-TTS and Qwen3-ASR?

They are hosted voice models from Nari Labs — the team behind the open-source Dia TTS model — built on Alibaba's Qwen3 speech models and served on Nari's low-latency inference stack. Qwen3-TTS 1.7B is text-to-speech; Qwen3-ASR 1.7B is speech-to-text (transcription). Announced September 14, 2026 as topping Coval's public voice-AI benchmark.

How fast and accurate are they on the Coval benchmark?

As of the September 14, 2026 benchmark, Nari Qwen3-ASR Fast reported roughly 44 ms median time-to-final-segment and about 3.6% word error rate (#1 latency, #2 accuracy), and Nari Qwen3-TTS Fast reported roughly 63 ms time-to-first-audio and about 3.8% WER (#2 latency, #1 accuracy). Treat exact numbers as a benchmark-day snapshot and confirm on Coval's page.

How much do Nari Qwen3-TTS and Qwen3-ASR cost?

Nari lists TTS at $10 per 1M characters on the Fast tier and $5 per 1M on Standard, and ASR at $0.12 per hour on Fast and $0.06 per hour on Standard — well below premium voice APIs. Self-hosting Nari's Apache-2.0 Qwen3-TTS engine works out to roughly $2 per 1M characters at full utilization. Confirm current pricing on Nari's site.

Can these models publish content to social platforms?

No. Qwen3-TTS generates audio and Qwen3-ASR generates transcripts; neither writes posts, makes video or images, captions clips, or schedules and publishes anywhere. To turn a cheap voiceover or a transcript into finished, on-brand posts across platforms you pair them with a content engine like Kompozy, which generates the formats and publishes across the eight social platforms plus blog and email.

Related news

← All AI news · Get started →