Wispr's first proprietary speech-recognition model, built for real-time dictation in real-world audio — accurate through background noise, music, accents, and far-field or whispered speech, powering the Wispr Flow app.
Last verified · 2026-09-17 · by Moe Ameen
Canto is the speech-recognition model built by Wispr — the San Francisco company behind the Wispr Flow dictation app — detailed on September 17, 2026 after a preview alongside its August 2026 Series B ($280 million at a $2 billion valuation). It's Wispr's first proprietary model, trained on a foundation of millions of hours of speech and text, and it's built for one thing: real-time dictation that stays accurate in the conditions people actually record in.
That's the differentiator. Most speech models are trained and benchmarked on clean, read-aloud audio, so they degrade the moment there's background noise, music, traffic, an interruption, a heavy accent, or far-field and whispered, low-volume speech. Wispr trained Canto specifically for that mess and reports word error rates in the hardest conditions falling from over 30% to roughly 5–10%. It supports contextual vocabulary so names and domain terms come through, and Wispr describes a two-stage training method — supervised fine-tuning followed by reinforcement learning (Group Relative Policy Optimization) — that lets the model learn from real user corrections.
On its own tests, Wispr says Canto posted the lowest word error rate on real Wispr Flow dictations against Google, OpenAI, AssemblyAI, and Deepgram, led all real-time models on a harder challenge set, and tied for the lowest error rate on the public LibriSpeech benchmark. Treat those as company-reported until independent benchmarks land. Canto powers the Wispr Flow app rather than shipping as a clearly standalone API at launch; Flow has free and paid tiers, and pricing for the model itself wasn't broken out — confirm current availability on wisprflow.ai.
The clean framing for a creator: Canto is an ear, not a studio. It's very good at turning messy spoken audio into accurate text, and it stops there. It doesn't generate video, images, or carousels, doesn't caption or reframe media per platform, and doesn't schedule or publish. Turning what you dictate into finished, distributed content is a separate job.
Canto's real unlock for a creator is where it lets you capture: not a quiet booth, but a walk, a job site, a car, a busy event — the places ideas actually happen. It nails the transcript in that chaos. But a transcript from a noisy voice memo is still just words, and the reason you recorded on the move was to make content, not to collect text files. [Kompozy](/) is the layer that turns that field-captured transcript into finished posts. Talk through a script or a rough idea wherever you are, let Canto (via Wispr Flow) produce clean text despite the noise, drop it into Kompozy, and it generates a [Persona Short](/glossary/persona-shorts) or [Persona Frames](/glossary/persona-frames) video your face-locked avatar reads to camera — so a memo taped on a walk becomes a talking-head clip without you ever being on set — plus a brand-exact [Carousel](/glossary/hyperframes), [Quote Graphics](/glossary/output-buckets), a Blog Article, and an Email Newsletter, all held to one voice by the [Persona Brief](/glossary/persona-brief).
The payoff is that your recording conditions stop gating your output. Because Canto is built for far-field and low-volume speech, you can capture a whole week of ideas hands-free and messy; because Kompozy generates net-new media a transcriber can't — persona video, images, carousels — and publishes it across the eight social platforms plus blog and email on [Autopilot](/glossary/autopilot) behind a per-post review gate, each of those rough captures ships as polished, on-brand content. Canto removes the requirement to record clean; Kompozy removes everything between a captured idea and a published calendar. It pairs the same way with [Wispr Flow](/ai-tools/wispr-flow), the app Canto powers.
Canto is Wispr's first proprietary speech-recognition model, detailed on September 17, 2026 and previewed alongside its August 2026 Series B. It's built for real-time dictation in real-world audio — background noise, music, accents, and far-field or whispered speech — and it powers the Wispr Flow app.
Most speech models are tuned and benchmarked on clean, read-aloud audio and degrade in noise. Wispr trained Canto specifically for real-world conditions — background noise, music, accents, and far-field or whispered speech — and reports word error rates in the hardest conditions falling from over 30% to roughly 5–10%, using reinforcement learning on real user corrections. Those figures are company-reported.
At launch Canto powers the Wispr Flow app rather than shipping as a clearly standalone API, and Wispr didn't break out pricing for the model itself. Flow has free and paid tiers; confirm current availability and prices on wisprflow.ai.
No. Canto is a recognition model — it produces accurate text and stops there, with no media, captioning, reframing, scheduling, or publishing. To turn transcribed words into carousels, quote cards, persona video, blogs, and scheduled posts, pair it with a content engine like Kompozy, which accepts pasted transcripts and generates and publishes finished content across nine platforms.