Cactus Compute's open on-device speech-to-text model — a single 16.9 MB file that transcribes audio on a CPU, with nothing sent off the device.
Last verified · 2026-10-08 · by Moe Ameen
Whistle is an open speech-to-text model from Cactus Compute, released on October 2, 2026 and authored by Jakub Mroz and Henry Ndubuaku. Its defining trait is size: the whole model is a single 16.9 MB file that runs on a CPU with no GPU and no dependencies, on the same Needle engine Cactus uses for its small on-device language models. Whisper base, the model it benchmarks against, is roughly 145 MB by comparison.
It transcribes up to 30 seconds of 16 kHz mono audio in one pass and supports English, German, French, Spanish, Italian, Dutch, and Polish, detecting the language automatically unless you specify one. Beyond a plain transcript it returns word-level timestamps with probabilities, can output speech embeddings (one row per 80 ms frame) without decoding text, and can bias recognition toward keywords you supply for names or jargon. A silence check runs before decoding, so quiet clips return an empty transcript rather than a guessed sentence. Crucially, the audio is processed on the device and never sent off it.
Cactus ships the Needle engine prebuilt for a wide range of targets — desktop, mobile, wearable, browser, and embedded hardware — and publishes the weights on Hugging Face and the source on GitHub, though the launch page does not name a specific license. Cactus reports lower word error rates than Whisper base on several benchmarks and higher on a few; treat those figures as vendor-reported until independent testing confirms them.
The honest framing for a creator: Whistle is a developer's building block, not a transcription app. There is no upload-a-file interface, the 30-second window means long recordings have to be chunked in code, and it covers seven languages. It gives you accurate text from short spoken audio, locally and for free — what you do with that text is up to the tools you add on top.
Whistle solves the first problem in a spoken-content workflow, and it solves it privately: you capture a thought out loud — a voice memo, a short take to camera, a dictated idea — and Whistle turns it into clean text on your own device, with the audio never leaving it. What it hands back is a transcript, and a transcript is not content. That is the exact handoff into Kompozy. Drop the transcribed text in as a source and Kompozy treats it as raw material for generation, fanning that one spoken idea into a Blog Article, an Email Newsletter, a Text Post or X thread, a brand-exact Carousel, and Quote Graphics — every piece written in one voice through your Persona Brief, then captioned, reframed per platform, and scheduled.
The pairing is strongest when privacy or offline capture matters. Dictate on a flight or in the field with no signal, let Whistle transcribe locally, and the only thing that ever touches the cloud is the finished, on-brand post Kompozy publishes — not the raw audio. And if the spoken source is video rather than a memo, you skip the manual chunking entirely: Kompozy transcribes the footage itself as part of clipping it into captioned vertical shorts. So Whistle covers the private, short, local text-capture case; Kompozy covers the long-form video case and everything downstream of the transcript.
An open on-device speech-to-text model released October 2, 2026. The whole model is a 16.9 MB file that runs on a CPU with no dependencies, transcribes up to 30 seconds of 16 kHz mono audio per pass across seven European languages, returns word-level timestamps, and keeps audio on the device.
Not in a single step — it processes up to 30 seconds per pass, so long audio has to be chunked in code. For long-form spoken video, Kompozy transcribes the footage itself while clipping it into captioned shorts, so you do not chunk anything by hand.
The weights and source are published openly, so there is no API or license fee; you pay only for compute, which is minimal on a CPU. The launch page does not name a specific license, so check the Hugging Face and GitHub pages before commercial use.
Feed the transcribed text into Kompozy as a source. It generates a blog, newsletter, X thread, carousel, and quote cards from the one idea in your brand voice, then captions, reframes, and schedules them across the eight social platforms plus blog and email.