// AI TOOLS · CANTO

Canto

Wispr's first proprietary speech-recognition model, built for real-time dictation in real-world audio — accurate through background noise, music, accents, and far-field or whispered speech, powering the Wispr Flow app.

Last verified · 2026-09-17 · by Moe Ameen

What Canto is

Canto is the speech-recognition model built by Wispr — the San Francisco company behind the Wispr Flow dictation app — detailed on September 17, 2026 after a preview alongside its August 2026 Series B ($280 million at a $2 billion valuation). It's Wispr's first proprietary model, trained on a foundation of millions of hours of speech and text, and it's built for one thing: real-time dictation that stays accurate in the conditions people actually record in.

That's the differentiator. Most speech models are trained and benchmarked on clean, read-aloud audio, so they degrade the moment there's background noise, music, traffic, an interruption, a heavy accent, or far-field and whispered, low-volume speech. Wispr trained Canto specifically for that mess and reports word error rates in the hardest conditions falling from over 30% to roughly 5–10%. It supports contextual vocabulary so names and domain terms come through, and Wispr describes a two-stage training method — supervised fine-tuning followed by reinforcement learning (Group Relative Policy Optimization) — that lets the model learn from real user corrections.

On its own tests, Wispr says Canto posted the lowest word error rate on real Wispr Flow dictations against Google, OpenAI, AssemblyAI, and Deepgram, led all real-time models on a harder challenge set, and tied for the lowest error rate on the public LibriSpeech benchmark. Treat those as company-reported until independent benchmarks land. Canto powers the Wispr Flow app rather than shipping as a clearly standalone API at launch; Flow has free and paid tiers, and pricing for the model itself wasn't broken out — confirm current availability on wisprflow.ai.

The clean framing for a creator: Canto is an ear, not a studio. It's very good at turning messy spoken audio into accurate text, and it stops there. It doesn't generate video, images, or carousels, doesn't caption or reframe media per platform, and doesn't schedule or publish. Turning what you dictate into finished, distributed content is a separate job.

What you can make with it

  • Accurate transcripts of audio recorded in real-world conditions — noise, music, traffic, and crowds
  • Clean voice-to-text of accented, far-field, or low-volume and whispered speech that other models garble
  • Real-time dictation into apps (via Wispr Flow) with names and jargon preserved through contextual vocabulary
  • Low-cleanup first drafts of scripts, hooks, captions, and outlines spoken aloud instead of typed
  • A reliable spoken-idea capture layer for anywhere you can't sit at a quiet desk
  • Note: outputs are text — not captions, video, or cross-platform posts

How Kompozy turns Canto output into content

Canto's real unlock for a creator is where it lets you capture: not a quiet booth, but a walk, a job site, a car, a busy event — the places ideas actually happen. It nails the transcript in that chaos. But a transcript from a noisy voice memo is still just words, and the reason you recorded on the move was to make content, not to collect text files. [Kompozy](/) is the layer that turns that field-captured transcript into finished posts. Talk through a script or a rough idea wherever you are, let Canto (via Wispr Flow) produce clean text despite the noise, drop it into Kompozy, and it generates a [Persona Short](/glossary/persona-shorts) or [Persona Frames](/glossary/persona-frames) video your face-locked avatar reads to camera — so a memo taped on a walk becomes a talking-head clip without you ever being on set — plus a brand-exact [Carousel](/glossary/hyperframes), [Quote Graphics](/glossary/output-buckets), a Blog Article, and an Email Newsletter, all held to one voice by the [Persona Brief](/glossary/persona-brief).

The payoff is that your recording conditions stop gating your output. Because Canto is built for far-field and low-volume speech, you can capture a whole week of ideas hands-free and messy; because Kompozy generates net-new media a transcriber can't — persona video, images, carousels — and publishes it across the eight social platforms plus blog and email on [Autopilot](/glossary/autopilot) behind a per-post review gate, each of those rough captures ships as polished, on-brand content. Canto removes the requirement to record clean; Kompozy removes everything between a captured idea and a published calendar. It pairs the same way with [Wispr Flow](/ai-tools/wispr-flow), the app Canto powers.

  1. Dictate your idea, script, or outline wherever you are — on a walk, in the car, on set — using Wispr Flow (powered by Canto), so noise doesn't wreck the transcript.
  2. Copy the clean text and paste it into Kompozy as your source.
  3. Pick your formats — Persona Short, Carousel, Quote Graphics, Blog Article, Text Posts — and let Kompozy generate them in one on-brand voice via the Persona Brief.
  4. Review the drafts in the per-post pipeline; tweak any caption, slide, or line inline.
  5. Schedule the set across the eight social platforms plus blog and email, and let Autopilot publish on cadence.

Frequently asked questions

What is Canto?

Canto is Wispr's first proprietary speech-recognition model, detailed on September 17, 2026 and previewed alongside its August 2026 Series B. It's built for real-time dictation in real-world audio — background noise, music, accents, and far-field or whispered speech — and it powers the Wispr Flow app.

How is Canto different from other speech models?

Most speech models are tuned and benchmarked on clean, read-aloud audio and degrade in noise. Wispr trained Canto specifically for real-world conditions — background noise, music, accents, and far-field or whispered speech — and reports word error rates in the hardest conditions falling from over 30% to roughly 5–10%, using reinforcement learning on real user corrections. Those figures are company-reported.

Is Canto available on its own, and what does it cost?

At launch Canto powers the Wispr Flow app rather than shipping as a clearly standalone API, and Wispr didn't break out pricing for the model itself. Flow has free and paid tiers; confirm current availability and prices on wisprflow.ai.

Can Canto turn my dictation into social posts or video?

No. Canto is a recognition model — it produces accurate text and stops there, with no media, captioning, reframing, scheduling, or publishing. To turn transcribed words into carousels, quote cards, persona video, blogs, and scheduled posts, pair it with a content engine like Kompozy, which accepts pasted transcripts and generates and publishes finished content across nine platforms.

Related tools

  • Wispr FlowAn AI dictation app that turns natural speech into clean, formatted text in any application — email, docs, chat, code editors — trimming filler and fixing formatting as you talk, across Mac, Windows, iOS, and Android.
  • YapA free, open-source macOS voice dictation app that transcribes on-device with Apple's Speech framework — press a shortcut, talk, and your words land in whatever field you were typing in, with no cloud, no account, and no model to download.
  • HyprnoteA privacy-first, open-source macOS notepad that transcribes and summarizes your meetings entirely on-device — no meeting bots, no cloud APIs, and no audio ever leaving your Mac.
  • DescriptAn AI audio and video editor you drive by editing the transcript — with the Underlord AI assistant, Studio Sound, AI avatars and voices, and translation, built for podcasts and video.

← All AI tools · Get started →