// AI TOOLS · NARI QWEN3-TTS & QWEN3-ASR

Nari Qwen3-TTS & Qwen3-ASR

Nari Labs' low-latency, low-cost voice models — Qwen3-TTS 1.7B for text-to-speech and Qwen3-ASR 1.7B for transcription — that topped Coval's public voice-AI benchmark in September 2026.

Last verified · 2026-09-14 · by Moe Ameen

What Nari Qwen3-TTS & Qwen3-ASR is

Nari Qwen3-TTS and Qwen3-ASR are two hosted voice models from Nari Labs, the team behind Dia — an open-source dialogue text-to-speech model released under Apache 2.0 and covered as a challenger to incumbents like ElevenLabs. Each is an optimized serving of one of Alibaba's Qwen3 speech models on Nari's own low-latency inference stack. Qwen3-TTS 1.7B converts text into natural, streaming speech; Qwen3-ASR 1.7B transcribes spoken audio into text. Both run on a "Fast" tier and a cheaper "Standard" tier.

On September 14, 2026, Nari announced that both models topped Coval's public voice-AI benchmark — a reproducible, open-source benchmark of TTS and STT providers published at benchmarks.coval.ai. Qwen3-ASR Fast placed first in latency (about 44 ms median time-to-final-segment) and second in accuracy (about 3.6% word error rate); Qwen3-TTS Fast placed second in latency (about 63 ms time-to-first-audio) and first in accuracy (about 3.8% WER). For reference, the benchmark measured the official Qwen3 TTS Flash realtime endpoint at 8.8% WER and 692 ms time-to-first-audio, so Nari's serving is markedly more accurate and far faster to first sound.

The engineering behind that is public: in August 2026 Nari open-sourced, under Apache 2.0, its Qwen3-TTS 1.7B inference engine, which reached sub-50 ms p95 time-to-first-audio at 10 requests per second on a single NVIDIA H100 — roughly $2 per 1M characters if you self-host at full utilization. The hosted API is priced aggressively: TTS at $10 per 1M characters (Fast) and $5 (Standard), ASR at $0.12 per hour (Fast) and $0.06 (Standard) — well under premium voice APIs. It exposes streaming and non-streaming generation over HTTP plus WebSocket input streaming for incremental text.

The honest framing: these are excellent, cheap audio and speech models, not a content tool. They turn text into a voice track or a recording into a transcript, and they do it fast and inexpensively — but they write no copy, make no video or images, caption nothing, and publish nowhere. Voices, language coverage, and any cloning inherit from the Qwen3 speech base, so confirm those specifics on Nari's docs. Treat exact latency, WER, and pricing figures as a benchmark-day snapshot.

What you can make with it

  • Cheap, natural voiceover from text via Qwen3-TTS — narration for faceless video, explainers, and audio versions of articles
  • Low-latency streaming speech for real-time voice agents and live assistants (sub-70 ms to first sound)
  • Clean transcripts of podcasts, webinars, calls, and voice memos via Qwen3-ASR at ~$0.06–$0.12 per hour
  • Bulk transcription of a back catalog of recordings into searchable, repurposable text
  • Self-hosted, near-marginal-cost TTS at scale using the Apache-2.0 Qwen3-TTS inference engine on your own H100
  • Multilingual speech and transcription drawing on the Qwen3 speech base (confirm current language coverage on Nari's docs)

How Kompozy turns Nari Qwen3-TTS & Qwen3-ASR output into content

Think of Nari as two cheap taps on the same audio layer — one that turns text into a voice track, one that turns a recording into text — and think of [Kompozy](/) as the engine that turns whatever comes out of either tap into finished, on-brand, published content. That division is clean because the tools never overlap: Nari does the audio round-trip fast and for almost nothing; Kompozy does the generation and the distribution Nari can't touch. It is a full AI content generation and multi-platform publishing engine, not a voice API.

The most valuable pairing runs the ASR tap first. You already speak more content than you'll ever type — a podcast, a client call, a walk-and-talk voice note — and at $0.06–$0.12 an hour, transcribing all of it is basically free. Feed that Qwen3-ASR transcript into Kompozy's Quick Ingest and one recording fans out under a [Persona Brief](/glossary/persona-brief) into [Clipped Shorts](/glossary/clipped-short), brand-exact Carousel Posts and Quote Graphics via [HyperFrames](/glossary/hyperframes), native Text Posts, a Blog Article, and an Email Newsletter — then publishes the set across the eight social platforms plus blog and email with [Autopilot](/glossary/autopilot). Run the TTS tap the other direction when you want a standalone narration track for a faceless Listicle Video or an audio version of a post — though note Kompozy's avatar formats like [Persona Shorts](/glossary/persona-shorts) already speak with HeyGen's native TTS, so you reach for Qwen3-TTS only when you specifically want a separate, dirt-cheap voice file. Nari owns the audio in and the audio out; Kompozy owns everything in between and after.

  1. Record once — a podcast, a livestream, a voice memo — and run it through Qwen3-ASR to get a clean transcript for a few cents.
  2. Drop the transcript into Kompozy's Quick Ingest as a source; the Persona Brief fixes your voice and angle across everything that follows.
  3. Generate the full set from that one source: Clipped Shorts, Carousel Posts, Quote Graphics, Text Posts, a Blog Article, and a Newsletter.
  4. When you want a standalone narration track, generate the script in Kompozy and voice it with low-cost Qwen3-TTS for a faceless or audio-version output.
  5. Review each piece behind the per-post gate, then schedule and publish across the eight social platforms plus blog and email with Autopilot.

Frequently asked questions

What are Nari Qwen3-TTS and Qwen3-ASR?

They are two hosted voice models from Nari Labs, the team behind the open-source Dia TTS model. Each is an optimized serving of one of Alibaba's Qwen3 speech models on Nari's low-latency stack: Qwen3-TTS 1.7B for text-to-speech and Qwen3-ASR 1.7B for transcription. Nari announced on September 14, 2026 that both topped Coval's public voice-AI benchmark.

How much do they cost?

Nari lists text-to-speech at $10 per 1M characters (Fast) and $5 per 1M (Standard), and speech-to-text at $0.12 per hour (Fast) and $0.06 per hour (Standard) — well below premium voice APIs. Self-hosting Nari's Apache-2.0 Qwen3-TTS engine works out to roughly $2 per 1M characters at full utilization on one H100. Confirm current pricing on Nari's site.

Are these better than Whisper or ElevenLabs?

On the September 2026 Coval benchmark, Nari led on TTS accuracy and STT latency and is far cheaper per unit than premium APIs. Whisper is the common open ASR baseline; ElevenLabs is a broader audio suite with cloning and licensed music. For the cheapest fast voice or transcription, Nari is compelling; for a full toolkit, the incumbents cover more. Benchmark on your own audio.

Can Nari Qwen3-TTS or Qwen3-ASR post content to social media?

No. They generate audio (TTS) or transcripts (ASR) only — they write no posts, make no video or images, caption nothing, and publish nowhere. To turn a cheap voiceover or a transcript into finished, on-brand posts across platforms you bring it into a content engine like Kompozy.

How do I use Nari's models with Kompozy?

The most useful path is to transcribe what you record with cheap Qwen3-ASR, then drop the transcript into Kompozy to fan one recording into clips, carousels, quote cards, text posts, a blog, and a newsletter, published across the eight social platforms plus blog and email. Use Qwen3-TTS when you want a standalone narration track; Kompozy's avatar formats already include native TTS.

Related tools

  • Qwen ScribeA free, open-source macOS app that runs Alibaba's Qwen3-ASR speech model entirely on Apple Silicon — private file transcription plus system-wide dictation, with no cloud, account, or API key.
  • Kokoro TTSAn open-weight, 82-million-parameter text-to-speech model that runs high-quality narration locally on a CPU — free, offline, and Apache-2.0 licensed for commercial use.
  • Whisper on Cloudflare Workers AIOpenAI's open-source Whisper speech-to-text model, served on Cloudflare's edge with a free daily allowance and per-audio-minute pricing.
  • Transcribe.cppOpen-source C/C++ library that runs 16+ speech-to-text model families locally on your own GPU via the ggml runtime — accurate, offline transcription with no per-minute API bill.
  • HeyGenAI avatar video platform that turns a text script into a talking-head video — in 175+ languages.

← All AI tools · Get started →