Cue review 2026: honest scoring on the voice-activated desktop agent, its local Gemma dictation, agentic tasks, latency, pricing, platform support, and who it actually fits.
Cue is one of the more genuinely useful voice tools on the desktop in 2026: fast, private, hands-free dictation into any app — powered by a local Gemma model via Ollama — wrapped around a real agent that reads your screen and runs multi-step tasks by voice. As a voice-input and desktop-automation tool it earns its praise. Its limits are scope and maturity: it produces text and actions, not finished content, it publishes nothing, and its paid tiers weren't priced yet. Excellent for talking instead of typing; not a content pipeline.
Cue is a voice-activated AI agent that lives on your Mac or Windows desktop, from the company Cue (founder and CEO Eli Li) at heycue.io. The pitch is simple and appealing: press a hotkey, speak, and it acts — dictating cleaned text at your cursor in any app, or running a multi-step task by reading what's on your screen and choosing the right tools. This review scores it as what it is, a voice-input and desktop-automation tool, not a content-creation suite, because grading it against the wrong job would be unfair.
The technical detail that put Cue in the news is on-device dictation. In a writeup on Google DeepMind's Gemmaverse, Cue described moving its text-polishing step to Google's open Gemma model — Gemma 4 E4B — running locally via Ollama: median latency fell from 876 ms to 488 ms, a 44% reduction on Apple Silicon, per-user dictation rose about 30% afterward, and the marginal cost of that step dropped to zero. A cloud model stays in place as a fallback when Ollama isn't running, so dictation keeps working either way.
Around the dictation sits the agent. Cue can read the screen and execute commands like drafting an email, analyzing on-screen data, or generating a document from a spoken brief. It runs on macOS 13+ (Apple Silicon and Intel) and Windows 10 build 1809+ (64-bit; Windows on ARM isn't supported yet). It's free to start with unlimited dictation and a daily allotment of agent credits; paid power-user tiers were described as coming but weren't priced at the time of writing.
I score it on dimensions that fit a voice desktop agent: dictation quality, speed and latency, agentic capability, on-device and privacy, ease of use, platform support, and value — plus, honestly, content and publishing reach, where it scores low because it makes text and actions and doesn't produce or publish content. Everything below reflects Cue's public state as of 2026-07-21; the product is young and paid tiers were still unannounced, so confirm current features, pricing, and access at heycue.io before relying on them.
Cue is a native desktop app for macOS and Windows that turns voice into text and actions. Its dictation layer lets you hold Option on a Mac (Alt on Windows) and speak; cleaned text with context-aware punctuation appears at your cursor in whatever app is focused, and the formatting adapts to that app (a chat message versus an email versus a terminal command). Its agent layer goes further: given a spoken command, it reads the screen, picks the right tools, and executes a multi-step task. Under the hood, the dictation's text-polishing runs Google's Gemma 4 E4B model locally via Ollama, with a cloud model as a fallback. What Cue does not do is anything on the content-production side. It doesn't generate carousels, blogs, newsletters, quote graphics, or video; it has no brand-voice or persona layer to keep output consistent across formats; it doesn't reframe or resize media per platform; and it doesn't schedule or publish to any channel. It is an input and automation tool, positioned to make the keyboard optional — not a studio for producing and distributing content.
Cue fits people who spend their day inside desktop apps and want a faster, hands-free way to get words onto the screen and to fire off routine actions — writers drafting into any editor, professionals clearing email and notes by voice, developers dictating into a terminal or IDE, and anyone who prefers talking to typing. Its local-model dictation makes it a strong pick where speed and privacy matter. It is not aimed at social-media creators who need finished, on-brand posts and a publishing calendar; for that job Cue is a capture front-end at best, and a dedicated content engine does the actual work.
| Dimension | Score | Why |
|---|---|---|
| Dictation quality | 4.5 / 5 | Fast, accurate, hands-free dictation into any app with context-aware punctuation and app-matched formatting. |
| Speed & latency | 4.5 / 5 | Running Gemma 4 E4B locally via Ollama cut median latency 44% (876 ms to 488 ms on Apple Silicon). |
| Agentic capability | 3.8 / 5 | Real multi-step desktop tasks by voice, but newer, credit-limited on the free tier, and reliant on cloud processing. |
| On-device & privacy | 4.3 / 5 | Local dictation model plus a stated policy of sending voice/screen context only at request time and not training on it. |
| Ease of use | 4.4 / 5 | Hotkey-driven, zero configuration, works in any app out of the box. |
| Platform support | 3.8 / 5 | Native Mac and Windows apps, but Windows on ARM isn't supported yet. |
| Value / pricing | 4.0 / 5 | Generous free tier (unlimited dictation + daily credits); paid power-user tiers were still unpriced at review time. |
| Content & publishing reach | 1.5 / 5 | Out of scope by design — it makes text and actions, not finished content, and publishes nothing. |
| Maturity & track record | 3.4 / 5 | A young product with a strong technical story but a short history and unshipped paid tiers. |
Cue's pricing, as of 2026-07-21, is unusually generous at the free end: unlimited voice dictation plus a daily allotment of agent credits, no credit card required. That generosity is only possible because the dictation model runs on-device via Ollama, so the marginal inference cost of that step is effectively zero — Cue can give away the part that would otherwise be expensive to serve. It's a smart alignment of technical architecture and business model.
The gap is at the top. Cue has described paid tiers for power users as coming, but had not published prices at review time, so anyone budgeting for heavy agent use is working with an unknown. The agentic tasks — which do rely on cloud processing — are the part most likely to be metered, and the daily free credit allotment hints that heavier automation will carry a cost. Until those tiers ship with numbers, the honest read is "free tier is excellent; paid economics are unproven."
For its category, that's a fair position. Cue is competing with dictation tools and voice assistants, and a strong free tier is the right way to win that market. Just don't confuse Cue's pricing with a content tool's: it meters voice input and desktop actions, not content generation or publishing, so the value math is entirely different from an engine you'd pay to produce and distribute posts.
| Use case | Fit | Why |
|---|---|---|
| Hands-free dictation into any desktop app | Strong | Exactly what Cue is built for — fast, private, and accurate across applications. |
| Voice-driven desktop automation (email, docs, on-screen tasks) | Strong | The agent reads your screen and executes multi-step commands from speech. |
| Capturing rough ideas and outlines by voice | Strong | A great front-end for getting messy ideas out of your head quickly. |
| Turning one idea into multi-format social content | Weak | Cue produces text and actions, not carousels, blogs, or video; it stops at the raw input. |
| Rendering avatar or persona video | Weak | Cue makes no video; you'd need a dedicated generation tool. |
| Scheduling and publishing across platforms | Weak | Cue has no publishing pipeline; it posts nothing to any channel. |
| Holding a consistent brand voice across content | Weak | Cue adapts formatting to the active app but has no brand-voice or persona layer. |
To be fair to both tools, Cue and Kompozy aren't really competitors — they sit at opposite ends of the same workflow, and the honest thing a reviewer can say is "know which end you're buying." Cue is capture and desktop automation: it's the fastest way to get an idea out of your head by voice, and it's very good at that. It is not, and doesn't claim to be, a content-production tool — it makes text and actions and then hands off.
Kompozy is the other end. It's a content generation and publishing engine: it takes an idea and, governed by a Persona Brief so the voice stays consistent, produces the finished formats — text posts, blog articles, brand-exact carousels via HyperFrames, photo posts, quote graphics, email newsletters, and net-new video like Persona Shorts and HeyGen avatar clips — then schedules and publishes across nine platforms plus email and blog. If you tried Cue hoping it would produce and post your content and found it stops at the dictation, that's not a Cue defect; it's a category mismatch, and Kompozy is the tool built for the job Cue leaves undone. Used together, Cue captures and Kompozy produces.
If you want fast, private, hands-free dictation and voice-driven desktop automation, yes — Cue is one of the stronger voice tools on the desktop in 2026, and its free tier is generous. If you're looking for a tool that produces and publishes social content, no — that's not what Cue does. It makes text and actions, not finished posts or video.
Cue's dictation runs Google's open Gemma model — Gemma 4 E4B — locally on your machine via Ollama, which it credits for cutting median latency from 876 ms to 488 ms on Apple Silicon. A cloud model is used as a fallback when Ollama isn't running.
Cue is a native desktop app for macOS 13 Ventura or later (Apple Silicon and Intel) and Windows 10 build 1809 or later (64-bit). Windows on ARM is not supported yet.
As of 2026-07-21, Cue is free to start with unlimited voice dictation and a daily allotment of agent credits, and no credit card. Paid tiers for power users were described as coming but hadn't been priced. Confirm current pricing at heycue.io.
No. Cue turns your speech into text and runs desktop actions; it doesn't generate carousels, blogs, newsletters, or video, and it has no scheduler or multi-platform publishing. To produce and post content, you'd pair it with a content engine like Kompozy.
Cue runs its dictation model locally via Ollama and states that voice and screen context are sent to providers only at request time to fulfill a task, not used to train models, and that it keeps local copies of your data. The agentic features do rely on cloud processing when a task requires it.
They do different jobs. Use Cue for hands-free dictation and voice-driven desktop automation; use Kompozy to generate finished multi-format content and publish it across platforms. Many creators use both — Cue to capture ideas by voice, Kompozy to turn those ideas into published posts.