Wispr detailed Canto, its first proprietary speech-recognition model, on September 17, 2026 — trained to transcribe through background noise, music, accents, and far-field or whispered speech, after previewing it alongside its August Series B.
2026-09-17 · by Moe Ameen
On September 17, 2026, Wispr — the San Francisco company behind the Wispr Flow dictation app — detailed Canto, its first proprietary speech-recognition model, on a dedicated launch page. The company previewed Canto in August 2026 alongside a $280 million Series B that valued it at $2 billion. Canto is built for real-time dictation and is trained on a foundation of millions of hours of speech and text.
The pitch is accuracy in the conditions people actually dictate in, not a quiet booth. Wispr says Canto is trained to hold up through background noise, music, traffic, interruptions, heavy accents, and far-field or whispered, low-volume speech. In the hardest conditions, the company reports word error rates falling from more than 30% of words to roughly 5–10%. Technically, Wispr describes a two-stage approach: supervised fine-tuning followed by reinforcement learning (Group Relative Policy Optimization) so the model can learn from real user corrections rather than optimizing for aggregate error alone.
On benchmarks, Wispr says Canto posted the lowest word error rate on an evaluation of real Wispr Flow dictations against models from Google, OpenAI, AssemblyAI, and Deepgram. On a harder audio challenge set, the company reports Canto finishing second only to a large Google model it says isn't suited to real-time use — and first among real-time models — and it claims a tie for the lowest error rate on the public LibriSpeech benchmark.
These figures come from Wispr's own testing, so treat them as vendor-reported until independent benchmarks land. Canto powers the Wispr Flow product, which offers free and paid tiers; pricing specific to the model wasn't broken out. Confirm current availability and figures on wisprflow.ai. The durable point stands: the speech-to-text race is now competing on messy, real-world audio, not clean read-aloud text.
The most useful way to read Canto is honest about what it changes: it makes the words you capture cleaner, not the content you publish. A near-perfect transcript of a noisy voice memo is still a block of text — it isn't a caption sized for Reels, a quote card, a carousel, or a video your audience can watch. That handoff, from clean words to finished posts, is exactly what [Kompozy](/) is for. Dictate a rough script or a stream-of-consciousness idea wherever you happen to be — Canto's whole point is that the noise won't wreck it — then paste the transcript into Kompozy, and under one [Persona Brief](/glossary/persona-brief) that locks your voice and banned words it becomes a [Persona Short](/glossary/persona-shorts) your face-locked avatar reads to camera, a brand-exact [Carousel](/glossary/hyperframes), [Quote Graphics](/glossary/output-buckets), a Blog Article, and an Email Newsletter, then ships across the eight social platforms plus blog and email on [Autopilot](/glossary/autopilot).
There's a faster move for today, too: the launch itself is a post. Point Kompozy at this news and it becomes a blog explainer of what Canto changes, a carousel of the benchmark claims, quote cards of the sharpest numbers, and a batch of captioned shorts — a full multi-platform push while the story is fresh. Canto is a better microphone for the real world; Kompozy is the studio that turns whatever you capture into a published week. See our coverage of [Wispr's $280M Series B](/news/wispr-280m-series-b-beyond-dictation) for the money side of the same story.
Canto is Wispr's first proprietary speech-recognition model, detailed on September 17, 2026 and previewed alongside its August 2026 Series B. It's built for real-time dictation in real-world conditions — background noise, music, accents, and far-field or whispered speech — and it powers the Wispr Flow app.
Wispr reports that in the hardest audio conditions, word error rates drop from over 30% to roughly 5–10%, and that on its own tests Canto led models from Google, OpenAI, AssemblyAI, and Deepgram on real dictations and tied for the lowest error rate on the public LibriSpeech benchmark. These are company-reported figures, so treat them as vendor claims until independent benchmarks confirm them.
Canto powers the Wispr Flow product rather than shipping as a clearly standalone API at launch, and Wispr didn't break out pricing for the model itself. Flow offers free and paid tiers; confirm current availability and prices on wisprflow.ai.
Cleaner transcription from messy, real-world audio makes dictation a more reliable way to capture ideas anywhere. But it still only produces text — you need a separate engine to turn those words into captions, video, carousels, and scheduled posts. Kompozy is built for that step and accepts pasted transcripts as a source.