Nari Labs' low-latency, low-cost voice models — Qwen3-TTS 1.7B for text-to-speech and Qwen3-ASR 1.7B for transcription — that topped Coval's public voice-AI benchmark in September 2026.
Last verified · 2026-09-14 · by Moe Ameen
Nari Qwen3-TTS and Qwen3-ASR are two hosted voice models from Nari Labs, the team behind Dia — an open-source dialogue text-to-speech model released under Apache 2.0 and covered as a challenger to incumbents like ElevenLabs. Each is an optimized serving of one of Alibaba's Qwen3 speech models on Nari's own low-latency inference stack. Qwen3-TTS 1.7B converts text into natural, streaming speech; Qwen3-ASR 1.7B transcribes spoken audio into text. Both run on a "Fast" tier and a cheaper "Standard" tier.
On September 14, 2026, Nari announced that both models topped Coval's public voice-AI benchmark — a reproducible, open-source benchmark of TTS and STT providers published at benchmarks.coval.ai. Qwen3-ASR Fast placed first in latency (about 44 ms median time-to-final-segment) and second in accuracy (about 3.6% word error rate); Qwen3-TTS Fast placed second in latency (about 63 ms time-to-first-audio) and first in accuracy (about 3.8% WER). For reference, the benchmark measured the official Qwen3 TTS Flash realtime endpoint at 8.8% WER and 692 ms time-to-first-audio, so Nari's serving is markedly more accurate and far faster to first sound.
The engineering behind that is public: in August 2026 Nari open-sourced, under Apache 2.0, its Qwen3-TTS 1.7B inference engine, which reached sub-50 ms p95 time-to-first-audio at 10 requests per second on a single NVIDIA H100 — roughly $2 per 1M characters if you self-host at full utilization. The hosted API is priced aggressively: TTS at $10 per 1M characters (Fast) and $5 (Standard), ASR at $0.12 per hour (Fast) and $0.06 (Standard) — well under premium voice APIs. It exposes streaming and non-streaming generation over HTTP plus WebSocket input streaming for incremental text.
The honest framing: these are excellent, cheap audio and speech models, not a content tool. They turn text into a voice track or a recording into a transcript, and they do it fast and inexpensively — but they write no copy, make no video or images, caption nothing, and publish nowhere. Voices, language coverage, and any cloning inherit from the Qwen3 speech base, so confirm those specifics on Nari's docs. Treat exact latency, WER, and pricing figures as a benchmark-day snapshot.
Think of Nari as two cheap taps on the same audio layer — one that turns text into a voice track, one that turns a recording into text — and think of [Kompozy](/) as the engine that turns whatever comes out of either tap into finished, on-brand, published content. That division is clean because the tools never overlap: Nari does the audio round-trip fast and for almost nothing; Kompozy does the generation and the distribution Nari can't touch. It is a full AI content generation and multi-platform publishing engine, not a voice API.
The most valuable pairing runs the ASR tap first. You already speak more content than you'll ever type — a podcast, a client call, a walk-and-talk voice note — and at $0.06–$0.12 an hour, transcribing all of it is basically free. Feed that Qwen3-ASR transcript into Kompozy's Quick Ingest and one recording fans out under a [Persona Brief](/glossary/persona-brief) into [Clipped Shorts](/glossary/clipped-short), brand-exact Carousel Posts and Quote Graphics via [HyperFrames](/glossary/hyperframes), native Text Posts, a Blog Article, and an Email Newsletter — then publishes the set across the eight social platforms plus blog and email with [Autopilot](/glossary/autopilot). Run the TTS tap the other direction when you want a standalone narration track for a faceless Listicle Video or an audio version of a post — though note Kompozy's avatar formats like [Persona Shorts](/glossary/persona-shorts) already speak with HeyGen's native TTS, so you reach for Qwen3-TTS only when you specifically want a separate, dirt-cheap voice file. Nari owns the audio in and the audio out; Kompozy owns everything in between and after.
They are two hosted voice models from Nari Labs, the team behind the open-source Dia TTS model. Each is an optimized serving of one of Alibaba's Qwen3 speech models on Nari's low-latency stack: Qwen3-TTS 1.7B for text-to-speech and Qwen3-ASR 1.7B for transcription. Nari announced on September 14, 2026 that both topped Coval's public voice-AI benchmark.
Nari lists text-to-speech at $10 per 1M characters (Fast) and $5 per 1M (Standard), and speech-to-text at $0.12 per hour (Fast) and $0.06 per hour (Standard) — well below premium voice APIs. Self-hosting Nari's Apache-2.0 Qwen3-TTS engine works out to roughly $2 per 1M characters at full utilization on one H100. Confirm current pricing on Nari's site.
On the September 2026 Coval benchmark, Nari led on TTS accuracy and STT latency and is far cheaper per unit than premium APIs. Whisper is the common open ASR baseline; ElevenLabs is a broader audio suite with cloning and licensed music. For the cheapest fast voice or transcription, Nari is compelling; for a full toolkit, the incumbents cover more. Benchmark on your own audio.
No. They generate audio (TTS) or transcripts (ASR) only — they write no posts, make no video or images, caption nothing, and publish nowhere. To turn a cheap voiceover or a transcript into finished, on-brand posts across platforms you bring it into a content engine like Kompozy.
The most useful path is to transcribe what you record with cheap Qwen3-ASR, then drop the transcript into Kompozy to fan one recording into clips, carousels, quote cards, text posts, a blog, and a newsletter, published across the eight social platforms plus blog and email. Use Qwen3-TTS when you want a standalone narration track; Kompozy's avatar formats already include native TTS.