// AI TOOLS · WHISTLE

Whistle

Cactus Compute's open on-device speech-to-text model — a single 16.9 MB file that transcribes audio on a CPU, with nothing sent off the device.

Last verified · 2026-10-08 · by Moe Ameen

What Whistle is

Whistle is an open speech-to-text model from Cactus Compute, released on October 2, 2026 and authored by Jakub Mroz and Henry Ndubuaku. Its defining trait is size: the whole model is a single 16.9 MB file that runs on a CPU with no GPU and no dependencies, on the same Needle engine Cactus uses for its small on-device language models. Whisper base, the model it benchmarks against, is roughly 145 MB by comparison.

It transcribes up to 30 seconds of 16 kHz mono audio in one pass and supports English, German, French, Spanish, Italian, Dutch, and Polish, detecting the language automatically unless you specify one. Beyond a plain transcript it returns word-level timestamps with probabilities, can output speech embeddings (one row per 80 ms frame) without decoding text, and can bias recognition toward keywords you supply for names or jargon. A silence check runs before decoding, so quiet clips return an empty transcript rather than a guessed sentence. Crucially, the audio is processed on the device and never sent off it.

Cactus ships the Needle engine prebuilt for a wide range of targets — desktop, mobile, wearable, browser, and embedded hardware — and publishes the weights on Hugging Face and the source on GitHub, though the launch page does not name a specific license. Cactus reports lower word error rates than Whisper base on several benchmarks and higher on a few; treat those figures as vendor-reported until independent testing confirms them.

The honest framing for a creator: Whistle is a developer's building block, not a transcription app. There is no upload-a-file interface, the 30-second window means long recordings have to be chunked in code, and it covers seven languages. It gives you accurate text from short spoken audio, locally and for free — what you do with that text is up to the tools you add on top.

What you can make with it

  • A local transcript of a short spoken clip, voice memo, or dictated note, with no audio leaving your device
  • Word-level timestamps for building frame-accurate captions on short video
  • Keyword-biased recognition that gets names, brands, and jargon right
  • Speech embeddings for search, classification, or a voice feature inside an app
  • An offline transcription step for a privacy-sensitive or no-network product
  • A tiny on-device voice layer for wearables, robots, or embedded hardware

How Kompozy turns Whistle output into content

Whistle solves the first problem in a spoken-content workflow, and it solves it privately: you capture a thought out loud — a voice memo, a short take to camera, a dictated idea — and Whistle turns it into clean text on your own device, with the audio never leaving it. What it hands back is a transcript, and a transcript is not content. That is the exact handoff into Kompozy. Drop the transcribed text in as a source and Kompozy treats it as raw material for generation, fanning that one spoken idea into a Blog Article, an Email Newsletter, a Text Post or X thread, a brand-exact Carousel, and Quote Graphics — every piece written in one voice through your Persona Brief, then captioned, reframed per platform, and scheduled.

The pairing is strongest when privacy or offline capture matters. Dictate on a flight or in the field with no signal, let Whistle transcribe locally, and the only thing that ever touches the cloud is the finished, on-brand post Kompozy publishes — not the raw audio. And if the spoken source is video rather than a memo, you skip the manual chunking entirely: Kompozy transcribes the footage itself as part of clipping it into captioned vertical shorts. So Whistle covers the private, short, local text-capture case; Kompozy covers the long-form video case and everything downstream of the transcript.

  1. Capture a short spoken idea and transcribe it locally with Whistle — a voice memo, a dictated note, or a quick take — keeping the audio on your device.
  2. Set a Persona Brief in Kompozy so everything generated from the transcript shares one voice and banned-word list.
  3. Paste the transcript into Kompozy as a source and generate across formats — blog, newsletter, X thread, carousel, quote cards — from the one idea.
  4. Let Kompozy caption any video, reframe each format for its platform, and keep the design brand-exact.
  5. Schedule and publish the set across the eight social platforms plus blog and email from one queue, or hand it to autopilot.

Frequently asked questions

What is Whistle by Cactus Compute?

An open on-device speech-to-text model released October 2, 2026. The whole model is a 16.9 MB file that runs on a CPU with no dependencies, transcribes up to 30 seconds of 16 kHz mono audio per pass across seven European languages, returns word-level timestamps, and keeps audio on the device.

Can Whistle transcribe long videos or podcasts?

Not in a single step — it processes up to 30 seconds per pass, so long audio has to be chunked in code. For long-form spoken video, Kompozy transcribes the footage itself while clipping it into captioned shorts, so you do not chunk anything by hand.

Is Whistle free to use?

The weights and source are published openly, so there is no API or license fee; you pay only for compute, which is minimal on a CPU. The launch page does not name a specific license, so check the Hugging Face and GitHub pages before commercial use.

How do I turn a Whistle transcript into social posts?

Feed the transcribed text into Kompozy as a source. It generates a blog, newsletter, X thread, carousel, and quote cards from the one idea in your brand voice, then captions, reframes, and schedules them across the eight social platforms plus blog and email.

Related tools

  • Whisper on Cloudflare Workers AI — OpenAI's open-source Whisper speech-to-text model, served on Cloudflare's edge with a free daily allowance and per-audio-minute pricing.
  • Gemini 3.5 Transcribe — Google's 2026 speech-to-text model that turns raw audio into clean, formatted text — automatically removing filler words like "ums" and "ahs" and fixing self-corrections.
  • Apple SpeechAnalyzer — Apple's on-device speech-to-text framework — a new proprietary transcription model, introduced at WWDC 2025, that benchmarks against OpenAI's Whisper.
  • Grok Voice Transcribe 2.0 — xAI's updated speech-to-text model, released September 18, 2026 — batch and real-time streaming transcription that xAI says is twice as accurate as v1.0 at the same price, with speaker labels and word-level timestamps included at no extra cost.

← All AI tools · Get started →