// AI NEWS · MODEL RELEASE

Wispr Launches Canto, a Speech Model Built for Real-World, Noisy Audio

Wispr detailed Canto, its first proprietary speech-recognition model, on September 17, 2026 — trained to transcribe through background noise, music, accents, and far-field or whispered speech, after previewing it alongside its August Series B.

2026-09-17 · by Moe Ameen

What happened

On September 17, 2026, Wispr — the San Francisco company behind the Wispr Flow dictation app — detailed Canto, its first proprietary speech-recognition model, on a dedicated launch page. The company previewed Canto in August 2026 alongside a $280 million Series B that valued it at $2 billion. Canto is built for real-time dictation and is trained on a foundation of millions of hours of speech and text.

The pitch is accuracy in the conditions people actually dictate in, not a quiet booth. Wispr says Canto is trained to hold up through background noise, music, traffic, interruptions, heavy accents, and far-field or whispered, low-volume speech. In the hardest conditions, the company reports word error rates falling from more than 30% of words to roughly 5–10%. Technically, Wispr describes a two-stage approach: supervised fine-tuning followed by reinforcement learning (Group Relative Policy Optimization) so the model can learn from real user corrections rather than optimizing for aggregate error alone.

On benchmarks, Wispr says Canto posted the lowest word error rate on an evaluation of real Wispr Flow dictations against models from Google, OpenAI, AssemblyAI, and Deepgram. On a harder audio challenge set, the company reports Canto finishing second only to a large Google model it says isn't suited to real-time use — and first among real-time models — and it claims a tie for the lowest error rate on the public LibriSpeech benchmark.

These figures come from Wispr's own testing, so treat them as vendor-reported until independent benchmarks land. Canto powers the Wispr Flow product, which offers free and paid tiers; pricing specific to the model wasn't broken out. Confirm current availability and figures on wisprflow.ai. The durable point stands: the speech-to-text race is now competing on messy, real-world audio, not clean read-aloud text.

Why it matters for creators

  • Real-world audio is where creators actually record. A model tuned for noise, music, and far-field speech means you can dictate a script or capture an idea on a walk, in a car, or on set without the transcript turning to mush.
  • Accuracy lowers the cleanup tax. The less time you spend fixing a transcript, the more dictation actually saves — which is the whole reason to talk instead of type.
  • It's still text. A more accurate transcriber gives you cleaner words, not a caption, a clip, a carousel, or a scheduled post — the transcript is the start of the content job, not the end.
  • The headline benchmarks are the company's own. Lowest-error claims against Google, OpenAI, AssemblyAI, and Deepgram are Wispr-reported — useful signal, but worth confirming with independent tests before treating them as settled.
  • The category is maturing fast. A proprietary model trained on real dictations, not just public read-aloud corpora, signals that voice input is becoming reliable enough to build a real workflow around.

How to act on this with Kompozy

The most useful way to read Canto is honest about what it changes: it makes the words you capture cleaner, not the content you publish. A near-perfect transcript of a noisy voice memo is still a block of text — it isn't a caption sized for Reels, a quote card, a carousel, or a video your audience can watch. That handoff, from clean words to finished posts, is exactly what [Kompozy](/) is for. Dictate a rough script or a stream-of-consciousness idea wherever you happen to be — Canto's whole point is that the noise won't wreck it — then paste the transcript into Kompozy, and under one [Persona Brief](/glossary/persona-brief) that locks your voice and banned words it becomes a [Persona Short](/glossary/persona-shorts) your face-locked avatar reads to camera, a brand-exact [Carousel](/glossary/hyperframes), [Quote Graphics](/glossary/output-buckets), a Blog Article, and an Email Newsletter, then ships across the eight social platforms plus blog and email on [Autopilot](/glossary/autopilot).

There's a faster move for today, too: the launch itself is a post. Point Kompozy at this news and it becomes a blog explainer of what Canto changes, a carousel of the benchmark claims, quote cards of the sharpest numbers, and a batch of captioned shorts — a full multi-platform push while the story is fresh. Canto is a better microphone for the real world; Kompozy is the studio that turns whatever you capture into a published week. See our coverage of [Wispr's $280M Series B](/news/wispr-280m-series-b-beyond-dictation) for the money side of the same story.

Quick takeaways

  • Wispr detailed Canto, its first proprietary speech-recognition model, on September 17, 2026, after previewing it alongside its August 2026 Series B.
  • Canto is built for real-world dictation — trained to stay accurate through background noise, music, accents, and far-field or whispered speech.
  • Wispr reports word error rates in the hardest conditions falling from over 30% to roughly 5–10%, using a two-stage training method (supervised fine-tuning plus reinforcement learning).
  • On the company's own benchmarks, Canto posted the lowest error rate on real Wispr Flow dictations versus Google, OpenAI, AssemblyAI, and Deepgram, and tied for lowest on LibriSpeech — vendor-reported figures worth confirming independently.
  • A more accurate transcriber still only produces text; turning those words into captions, video, carousels, and scheduled posts needs a content engine like Kompozy.

Frequently asked questions

What is Canto?

Canto is Wispr's first proprietary speech-recognition model, detailed on September 17, 2026 and previewed alongside its August 2026 Series B. It's built for real-time dictation in real-world conditions — background noise, music, accents, and far-field or whispered speech — and it powers the Wispr Flow app.

How accurate is Canto?

Wispr reports that in the hardest audio conditions, word error rates drop from over 30% to roughly 5–10%, and that on its own tests Canto led models from Google, OpenAI, AssemblyAI, and Deepgram on real dictations and tied for the lowest error rate on the public LibriSpeech benchmark. These are company-reported figures, so treat them as vendor claims until independent benchmarks confirm them.

Is Canto available separately, and what does it cost?

Canto powers the Wispr Flow product rather than shipping as a clearly standalone API at launch, and Wispr didn't break out pricing for the model itself. Flow offers free and paid tiers; confirm current availability and prices on wisprflow.ai.

What does Canto mean for content creators?

Cleaner transcription from messy, real-world audio makes dictation a more reliable way to capture ideas anywhere. But it still only produces text — you need a separate engine to turn those words into captions, video, carousels, and scheduled posts. Kompozy is built for that step and accepts pasted transcripts as a source.

Related news

← All AI news · Get started →