A router for voice AI — one OpenAI-compatible API key that sends each session to the speech-to-text, LLM, and text-to-speech models benchmarked as best for your language, latency, cost, and quality.
Last verified · 2026-08-18 · by Moe Ameen
Speko (speko.ai) is a routing platform for voice AI — the "OpenRouter for voice." Instead of integrating ElevenLabs, OpenAI, Deepgram, AssemblyAI, Cartesia, and every other speech provider separately, you point your voice agent at one OpenAI-compatible endpoint and a single API key. Speko forwards each session to the speech-to-text, language, and text-to-speech models its own benchmarks say win for your language and use case; when the numbers change, the routing changes and you rewrite nothing. It launched publicly out of Y Combinator's Summer 2026 batch, founded by Beknazar Abdikamalov — previously co-founder and CTO of the voice-AI company Hupo — with a small San Francisco team.
The benchmark data is the core of the pitch. Speko measures dozens of speech and language models across roughly ten languages — English, Arabic, French, German, Hindi, Norwegian, Spanish, Tamil, and Telugu among them — on word error rate, finalize latency (time from the caller stopping to a finalized transcript), time to first token, and cost per minute, and publishes the results openly at benchmarks.speko.ai. The argument the data makes: English-only leaderboards mislead you, because no single model is best in every language — the top transcription model for Arabic is often not the top one for Spanish. Routing per language, rather than picking one vendor globally, is how you get accuracy everywhere.
Integration is meant to be a base-URL swap. Speko works with the common real-time voice stacks — LiveKit and Pipecat — and exposes an MCP connection for Claude and Cursor, so an existing agent can adopt it without a rewrite. You can also build a new voice agent from scratch on the platform with model selection handled automatically. Speko lists STT pricing that scales with accuracy (roughly a fraction of a cent up to a couple of cents per minute depending on the model's error rate) and cites SOC 2 Type II, HIPAA, and GDPR compliance. Confirm live model coverage, per-model rates, and language support on speko.ai and benchmarks.speko.ai, since the catalog and numbers move.
The honest boundary: Speko is developer infrastructure for building real-time voice applications — phone agents, transcription pipelines, voice assistants. It routes and benchmarks the models; it is not a content tool. It does not write posts in your brand voice, generate persona or avatar video, build carousels, reframe media per platform, or schedule and publish to a social feed. It gets a voice application the best speech model for the moment; turning your product or expertise into distributed content is a separate job.
Speko answers a developer's question — "which speech model should this session run on, in this language, right now." It has nothing to say about marketing the thing you built with it. And that is the gap most voice-AI founders and agencies fall into: they ship an impressive multilingual agent and then have no content engine to turn demos, docs, and product updates into an audience. [Kompozy](/) is that engine. Point it at a demo recording, a changelog, or a founder's talk and it produces a week of posts: a [Persona Short](/glossary/persona-shorts) where your face-locked avatar explains what the agent does, a brand-exact [Carousel](/glossary/hyperframes) of the benchmark story, [Clipped Shorts](/glossary/output-buckets) from a longer demo, Quote Graphics from the sharpest lines, plus a launch Blog Article and Email Newsletter — then reframes each to 9:16, 1:1, and 16:9 and publishes across the eight social platforms plus blog and email from one queue.
The neat part is that Kompozy already owns the voice layer you would otherwise wire Speko for: its persona video runs on HeyGen's native text-to-speech, so you get a talking, on-brand avatar without integrating a TTS provider yourself. Speko is where multilingual voice *applications* get the best model per language; Kompozy is where a brand's voice *content* gets generated and shipped, governed by a [Persona Brief](/glossary/persona-brief) and a per-post review gate before it goes out on [Autopilot](/glossary/autopilot). Build the agent on Speko; build the audience on Kompozy.
Speko is a routing platform for voice AI — the "OpenRouter for voice." One OpenAI-compatible API key sends each session to the speech-to-text, language, and text-to-speech models its benchmarks say win for your language, latency, cost, and quality. It launched out of Y Combinator's Summer 2026 batch.
Speko benchmarks dozens of speech and language models across roughly ten languages on word error rate, finalize latency, time to first token, and cost per minute, and publishes the results at benchmarks.speko.ai. It routes each session to the model that measures best for your language and rules; when the numbers change, the routing changes with no code edit.
It routes across many speech-to-text, text-to-speech, and speech-to-speech providers — including ElevenLabs, OpenAI, Deepgram, AssemblyAI, Cartesia, and Google — behind one key. It works with LiveKit and Pipecat and exposes an MCP connection for Claude and Cursor. Confirm the current model list on speko.ai.
No. Speko is developer infrastructure for building real-time voice applications — it routes and benchmarks speech models and stops there. It has no brand-voice writer, no video or carousel generation, and no scheduler. To turn a voice product or expertise into captioned, on-brand posts across platforms, use a content engine like Kompozy.
Build your voice agent on Speko, then bring a demo, transcript, or update into Kompozy. Kompozy spins it into a persona short, carousel, clipped shorts, quote cards, a blog, and a newsletter under your Persona Brief and publishes them across nine destinations. Kompozy's persona video already ships with native voice, so you can market a voice product without wiring a TTS provider yourself.