// ON-DEVICE SPEECH-TO-TEXT REVIEW

Whistle Review (2026): Is Cactus Compute's 16.9 MB On-Device Speech-to-Text Model Worth It?

Whistle review (2026): Cactus Compute's 16.9 MB on-device speech-to-text model, scored honestly on accuracy, the 30-second limit, and who it really fits.

Last verified · 2026-10-08 · by Moe Ameen
The verdict
3.9 / 5

Whistle is a genuinely impressive piece of engineering: a 16.9 MB speech-to-text model that runs entirely on a CPU, keeps audio on the device, and reports lower word error rates than the far larger Whisper base on several benchmarks. But it is a developer model, not a transcription app — it handles up to 30 seconds of 16 kHz mono audio per pass across seven European languages, so turning a podcast or a long video into a clean transcript means chunking the audio and writing code. Lean into it for on-device and embedded voice features; look elsewhere if you want drag-and-drop long-form transcription.

Cactus Compute released Whistle on October 2, 2026, and the headline number is the whole story: the entire model is a single 16.9 MB file that runs on a CPU with no GPU and no dependencies, on the same Needle engine Cactus uses for its small on-device language models. For comparison, Whisper base — the model Whistle benchmarks against — is roughly 145 MB. The pitch is speech recognition small enough to live on a phone, a wearable, or a microcontroller, with audio that never leaves the device.

This review scores Whistle for what it is and stays honest about what it is not. I build a content engine, so I will be exact about the line between a model and a product. Whistle is a model — weights on Hugging Face, source on GitHub, a CLI and a package install — not an app with a transcription interface you upload a file into.

The constraints decide the fit. Whistle transcribes up to 30 seconds of 16 kHz mono audio in a single pass and supports English, German, French, Spanish, Italian, Dutch, and Polish, auto-detecting the language unless you name one. It returns word-level timestamps, can output speech embeddings, and can bias recognition toward keywords you supply. Those are developer features, and good ones. But a creator with an hour-long interview does not get a transcript by dropping the file in; someone has to split the audio into chunks and stitch the results back together.

Everything below reconciles against Cactus's own launch page and its published benchmarks on 2026-10-08. Where a figure is Cactus-reported rather than independently tested, this review says so.

What Whistle (Cactus Compute) is

Whistle is an open speech recognition model from Cactus Compute, authored by Jakub Mroz and Henry Ndubuaku and released on October 2, 2026. Under the hood it takes log-mel audio features through a convolutional stem and eight encoder blocks built on Cactus's Simple Attention design, with an eight-layer decoder using gated cross-attention and a five-beam search over a vocabulary of 8,192 text pieces plus seven language tokens. A silence check runs before decoding, so quiet or steady-noise clips return an empty transcript instead of a hallucinated sentence. It runs on the Needle CPU engine Cactus ships prebuilt for a long list of targets — desktop, mobile, wearable, browser, and embedded — and the weights and source are published openly, though the launch page does not name a specific license. On accuracy, Cactus reports lower word error rates than Whisper base on several benchmarks (LibriSpeech, SPGISpeech, Earnings-22, and the FLEURS average) while trailing it on a few others (TED-LIUM, AMI, and the MLS average) — a strong result given Whistle is a fraction of the size. On an Apple M4 Pro CPU it reports time-to-first-token in the low tens of milliseconds and decoding well over a thousand tokens per second. Treat those headline numbers as vendor-reported until independent testing lands.

Who Whistle (Cactus Compute) is for

Whistle fits developers and teams building on-device voice features: a transcription step inside an app, a wearable or robotics command layer, a privacy-sensitive product that cannot send audio to the cloud, or an offline tool that has to work with no network. For short utterances — voice commands, dictated notes, clip-length captures — in one of its seven supported languages, it packs a remarkable amount of capability into 16.9 MB. It is a poor fit for a non-developer who wants to upload a podcast or a long video and get a transcript back: there is no app, the 30-second window means you chunk the audio yourself, and languages outside its European set are unsupported. For that use, a hosted transcription product or a content engine that transcribes for you is the better tool.

Scoring breakdown

DimensionScoreWhy
Transcription accuracy4.0 / 5Reports lower word error rates than Whisper base on several benchmarks and higher on a few; strong for its size, but the figures are vendor-reported.
Model size & footprint4.8 / 5A single 16.9 MB file — roughly a ninth of Whisper base — running on a CPU with no GPU or dependencies.
On-device speed / latency4.6 / 5Time-to-first-token in the low tens of milliseconds and decoding over a thousand tokens per second on an M4 Pro CPU.
Privacy (audio stays local)4.7 / 5All processing happens on the device; no audio is sent off it, which is the whole point for sensitive use.
Language coverage3.3 / 5Seven European languages with auto-detection; nothing outside that set, so broad multilingual needs are unmet.
Long-form / creator readiness2.2 / 5The 30-second, 16 kHz mono per-pass limit means long recordings must be chunked and stitched in code — not built for podcasts or full videos.
Developer experience & deployment4.2 / 5Prebuilt for a wide range of targets, no dependencies, weights on Hugging Face, source on GitHub — but it is code, not a product.
Extra features (timestamps, embeddings, keyword biasing)4.3 / 5Word-level timestamps, speech embeddings per 80 ms frame, and keyword biasing go beyond a plain transcript.
Out-of-the-box usability for non-developers1.8 / 5There is no GUI or app. For anyone who cannot integrate a model, it is effectively unusable on its own.
Value (open weights, free to run)4.5 / 5Open weights and source, so no API or license fee — you pay only for the minimal CPU compute.

Pros and cons

Pros

  • Runs entirely on-device — audio never leaves the machine, a real privacy and compliance win.
  • A single 16.9 MB file on a CPU with no GPU or dependencies, and it works offline.
  • Reports lower word error rates than the far larger Whisper base on several benchmarks.
  • Word-level timestamps, speech embeddings, and keyword biasing are built in.
  • Open weights and source, so there is no API or license fee — just compute.
  • Prebuilt for a wide range of targets, from desktop to wearables to microcontrollers.

Cons

  • Transcribes only up to 30 seconds of 16 kHz mono audio per pass — long files need manual chunking in code.
  • Supports seven European languages only; no coverage beyond that set.
  • It is a model, not an app — there is no interface, so a non-developer cannot use it on its own.
  • Produces a transcript and nothing downstream: no content, captions, design, or publishing.
  • Benchmark figures are Cactus-reported and await independent testing.
  • The launch page names no specific license, so commercial terms need checking before you rely on it.

Pricing analysis

Whistle is free in the sense that matters to a developer: the weights and source are published openly, so there is no license fee or per-minute API charge to run it. Your cost is compute — and because it runs on a CPU in a 16.9 MB footprint, that cost is close to nothing per transcription. For an on-device or embedded product that would otherwise pay a cloud STT API per minute at scale, that is the entire argument.

The catch is that "free model" is not "free product." Someone has to integrate it, handle the 30-second chunking for anything longer, build the interface, and maintain it. For a developer that is routine work; for a creator or a small team without engineering, the real cost is the build, not the model. A hosted transcription service charges per minute precisely because it absorbs that work for you.

So the fair way to price Whistle is against alternatives in its own lane — other on-device or open STT models — not against a creator transcription app. Against a cloud API, it can be dramatically cheaper at volume and far better for privacy. Against a finished product you can use today, it is not competing on the same axis: it trades money for engineering time.

Use-case fit

Use caseFitWhy
On-device voice commands or dictation inside an appStrongShort utterances, low latency, and no cloud are exactly the design target.
Privacy-sensitive transcription where audio must stay localStrongNothing leaves the device, which a cloud service cannot guarantee.
Embedded, wearable, or robotics speechStrongTiny CPU-only footprint, prebuilt for many targets, works offline.
Keyword-biased recognition for names and jargonStrongBuilt-in keyword biasing improves domain-specific terms.
Frame-accurate captions for short clipsOKWord-level timestamps are there, but you handle the chunking and the rendering yourself.
Transcribing a full podcast or long video in one stepWeakThe 30-second-per-pass limit forces manual chunking and stitching for anything long.
Transcribing languages outside its European setWeakOnly seven languages are supported.
Drag-and-drop transcription for a non-developerWeakIt is a model, not an app — there is no interface to upload a file into.

Alternatives worth considering

  • Whisper / whisper.cpp — the open baseline Whistle benchmarks against; larger, but broader language coverage and a bigger ecosystem
  • Apple SpeechAnalyzer — on-device transcription built into Apple platforms, if you only ship to Apple hardware
  • A hosted STT API — best when you want long-form transcription without building the pipeline
  • Kompozy — best if your real goal is turning spoken content into finished, published posts rather than just a transcript

How Kompozy compares

Kompozy is not a speech-to-text model and does not compete with Whistle on accuracy or footprint — if your job is embedding recognition in a device, Whistle is the right tool and Kompozy is not in the conversation. The comparison only matters if a transcript was never the actual goal. Most creators do not want a transcript; they want the posts that come out of one. Kompozy treats transcription as an internal step: it transcribes your long video or podcast as part of clipping it into vertical shorts, auto-captioning them, and repurposing the spoken content into a blog, a newsletter, an X thread, carousels, and quote graphics.

So the honest split is by where you sit. A developer building a voice feature wires Whistle into code and owns everything after the transcript. A creator who records a talk and wants it everywhere by tomorrow hands Kompozy the video and never sees a model, a 30-second chunk, or a line of code — the transcription happens on the way to finished, scheduled posts across the eight social platforms plus blog and email. Whistle gives you the text; Kompozy gives you the content the text was always meant to become.

Frequently asked questions

What is Whistle by Cactus Compute?

It is an open, on-device speech-to-text model released on October 2, 2026. The whole model is a single 16.9 MB file that runs on a CPU with no GPU or dependencies, transcribes up to 30 seconds of 16 kHz mono audio per pass across seven European languages, returns word-level timestamps, and keeps audio on the device.

Is Whistle more accurate than Whisper?

Cactus reports lower word error rates than Whisper base on several benchmarks (LibriSpeech, SPGISpeech, Earnings-22, and the FLEURS average) and higher on a few (TED-LIUM, AMI, and the MLS average), while being roughly a ninth of Whisper base's size. Those figures are vendor-reported, so treat them as a strong indication pending independent testing.

Can Whistle transcribe a long podcast or video?

Not in one step. It processes up to 30 seconds of audio per pass, so long files have to be split into chunks and the results stitched back together in code. It is built for short utterances and on-device use, not drag-and-drop long-form transcription.

What languages does Whistle support?

English, German, French, Spanish, Italian, Dutch, and Polish. It auto-detects the language unless you specify one, and languages outside that set are not supported.

Is Whistle free?

The weights and source are published openly, so there is no API or license fee — you pay only for the compute to run it, which is minimal on a CPU. The launch page does not name a specific license, so check the Hugging Face and GitHub pages before relying on it commercially.

Who should use Whistle?

Developers building on-device, embedded, offline, or privacy-sensitive voice features in a supported language. It is a model, not an app, so a non-developer who just wants to upload a file and get a transcript is better served by a hosted service.

How do I turn Whistle transcripts into content?

Use Whistle (or Kompozy's built-in transcription) to get the text, then feed the spoken content to Kompozy, which repurposes it into clipped shorts, captioned video, a blog, a newsletter, threads, carousels, and quote graphics, then schedules them across the eight social platforms plus blog and email.

Related deep guides

See Whistle (Cactus Compute) vs Kompozy comparison → · Get Started →