The Y Combinator Summer 2026 startup gives voice apps one OpenAI-compatible API key that routes speech-to-text, LLM, and text-to-speech to whichever models its open benchmarks rank best for a given language, latency, cost, and quality — no code change when the numbers move.
2026-08-18 · by Moe Ameen
Speko (speko.ai) launched publicly out of Y Combinator's Summer 2026 batch as a routing layer for voice AI — explicitly modeled on OpenRouter, but for speech instead of text. The pitch: instead of integrating ElevenLabs, OpenAI, Deepgram, AssemblyAI, Cartesia, and other providers one by one, a developer points a voice agent at a single OpenAI-compatible endpoint and key, and Speko forwards each session to the speech-to-text, language, and text-to-speech models its benchmarks measure as best for that language and use case. When the benchmark numbers change, the routing changes and the developer rewrites nothing.
The company was founded by Beknazar Abdikamalov, previously co-founder and CTO of the voice-AI firm Hupo, with a small San Francisco team. Its central claim is that English-only leaderboards mislead teams building for global users, because no single model is best in every language — the top transcription model for Arabic is often not the top one for Spanish or Tamil. Speko publishes its own language-by-language benchmarks openly at benchmarks.speko.ai, scoring dozens of speech and language models across roughly ten languages on word error rate, finalize latency, time to first token, and cost per minute.
On the integration side, Speko works with the common real-time voice stacks — LiveKit and Pipecat — and exposes an MCP connection so Claude or Cursor can reach it, positioning adoption as a base-URL swap rather than a rewrite. It lists usage-based speech-to-text pricing that scales with model accuracy and cites SOC 2 Type II, HIPAA, and GDPR compliance. As with any fast-moving launch, confirm the live model list, language coverage, and rates on speko.ai before building around them.
There are two ways to act on this launch, and they map to two different readers. If you are building a voice product, Speko is a genuine convenience — but the thing you will still lack is a way to turn that product into a content pipeline. If you are a creator covering the AI space, the move today is faster: this is a clean, timely subject to make content *about*. Either way the missing layer is the same one, and it is where [Kompozy](/) lives.
Take the launch itself as source material. Drop the "OpenRouter for voice" framing and Speko's language-by-language benchmark story into Kompozy and it generates the spread in minutes — a [Persona Short](/glossary/persona-shorts) where your face-locked avatar explains why no single voice model wins every language, a brand-exact [Carousel](/glossary/hyperframes) walking through the benchmark angle, [Quote Graphics](/glossary/output-buckets) from the sharpest lines, and a short explainer Blog Article and Email Newsletter — each written in one voice through your [Persona Brief](/glossary/persona-brief), reframed to 9:16, 1:1, and 16:9, and pushed across the eight social platforms plus blog and email from one queue on [Autopilot](/glossary/autopilot). Kompozy is model-agnostic by design (bring-your-own-key on the Founding tier), so the same "the model is swappable, the distribution is the moat" logic that makes Speko interesting is exactly why the publishing layer is the part worth owning. Speko routes the voice; Kompozy ships the story.
Speko is a routing platform for voice AI — the "OpenRouter for voice." One OpenAI-compatible API key routes each session to the speech-to-text, language, and text-to-speech models its benchmarks rank best for a given language, latency, cost, and quality. It launched out of Y Combinator's Summer 2026 batch.
Because no single speech model is best in every language — English-only leaderboards hide that. Speko benchmarks dozens of models across roughly ten languages and routes each session to the top performer for that language, so accuracy holds up outside English.
No. Speko is developer infrastructure for building real-time voice applications — it routes and benchmarks speech models and stops there. Turning a voice product or a launch like this into captioned, on-brand posts across platforms is a separate job, handled by a content engine like Kompozy.