A voice-activated AI agent that lives on your Mac or Windows desktop — press a hotkey, speak, and it dictates in any app or runs multi-step agentic tasks by reading your screen.
Last verified · 2026-07-21 · by Moe Ameen
Cue is a voice-activated AI agent that runs on your desktop rather than in a browser tab. You press a hotkey, speak, and it acts: at the simple end that's dictation — hold Option on a Mac (Alt on Windows) and your words appear at the cursor in whatever app is focused, with context-aware punctuation and formatting that adapts to where you're typing (a Slack message, an email, a terminal command). At the more ambitious end it's an agent that reads what's on your screen, picks the right tools, and executes multi-step tasks from a spoken command. It's built by Cue, whose founder and CEO is Eli Li, and it's distributed at heycue.io.
What makes Cue notable technically is that its dictation runs a Google Gemma model locally on your machine via Ollama, not in the cloud. In a writeup published on Google DeepMind's Gemmaverse, Cue described moving its text-polishing step to Gemma 4 E4B running on-device: median latency dropped from 876 ms to 488 ms — a 44% reduction on Apple Silicon — per-user dictation usage rose about 30% afterward, and the marginal inference cost of that step fell to zero. A cloud model stays in place as a fallback: if Ollama isn't running, Cue detects that at startup and routes to the cloud so dictation keeps working.
Cue runs on macOS 13 Ventura or later (Apple Silicon and Intel) and Windows 10 build 1809 or later (64-bit; Windows on ARM isn't supported yet). It's free to start, with unlimited voice dictation and a daily allotment of agent credits, and the company has described paid tiers for power users as coming. On privacy, Cue keeps local copies of your data and says voice and screen context are sent to providers only at request time to fulfill a task, not used to train models.
The honest framing: Cue is an input and automation layer — a fast, hands-free way to get words onto the screen and to fire off desktop actions by voice. It is not a content studio. It doesn't generate a carousel, render an avatar video, write and format a blog, or publish anything to a social platform. It turns speech into text and actions; what you do with that text downstream is a separate job.
The most useful way to think about Cue next to Kompozy is capture versus production. Cue is the fastest way to get a raw idea out of your head — you talk, and clean text lands wherever your cursor is. But dictated text is a starting point, not a finished post; a talked-through recap of your day, a spoken outline for a video, or a rough pitch is exactly the kind of source material Kompozy is built to turn into real, on-brand, multi-format content. Cue handles the front of the pipeline; Kompozy handles everything after.
Concretely: dictate the messy version of an idea with Cue — the story, the three points you want to make, the offer — into a note or straight into Kompozy's ingest. Kompozy then runs it through your Persona Brief (so the voice matches every time) and generates the finished formats that idea deserves: a Text Post and a Blog Article, a brand-exact Carousel via HyperFrames, Photo Posts and Quote Graphics, an Email Newsletter, and net-new video Cue can't touch — Persona Shorts and HeyGen avatar video with a face-locked recurring persona, plus Clipped Shorts from longer footage. Then Kompozy schedules and publishes the whole set across nine platforms (Instagram, Facebook, TikTok, YouTube, LinkedIn, X, Pinterest, Threads) plus email and blog destinations. Cue removes the keyboard from idea capture; Kompozy removes the manual work of turning that idea into a week of published content.
Cue is a voice-activated AI agent that runs on your Mac or Windows desktop. You press a hotkey and speak: it dictates into whatever app is focused, and it can also run multi-step agentic tasks by reading your screen and choosing the right tools. It's made by Cue (founder and CEO Eli Li) and distributed at heycue.io.
Cue's dictation runs a Google Gemma model — Gemma 4 E4B — locally on your machine via Ollama, which it credits for cutting median latency from 876 ms to 488 ms on Apple Silicon. A cloud model is used as a fallback when Ollama isn't running.
Cue is free to start with unlimited voice dictation and a daily allotment of agent credits; the company has said paid tiers for power users are coming. It runs on macOS 13 or later (Apple Silicon and Intel) and Windows 10 build 1809 or later (64-bit).
Not on its own. Cue turns your speech into text and fires desktop actions; it doesn't generate carousels, avatar videos, blogs, or newsletters, and it doesn't publish to social platforms. Dictate the idea with Cue, then use a content engine like Kompozy to generate the finished multi-format content and schedule it across platforms.
Different jobs. Cue is a voice input and desktop-automation layer — fast, hands-free capture and task execution. Kompozy is a content generation and publishing engine that takes an idea and produces posts, images, carousels, blogs, newsletters, and persona/avatar video, then publishes them across nine platforms. They pair well: Cue captures, Kompozy produces.