// AI TOOLS · SPEKO

Speko

A router for voice AI — one OpenAI-compatible API key that sends each session to the speech-to-text, LLM, and text-to-speech models benchmarked as best for your language, latency, cost, and quality.

Last verified · 2026-08-18 · by Moe Ameen

What Speko is

Speko (speko.ai) is a routing platform for voice AI — the "OpenRouter for voice." Instead of integrating ElevenLabs, OpenAI, Deepgram, AssemblyAI, Cartesia, and every other speech provider separately, you point your voice agent at one OpenAI-compatible endpoint and a single API key. Speko forwards each session to the speech-to-text, language, and text-to-speech models its own benchmarks say win for your language and use case; when the numbers change, the routing changes and you rewrite nothing. It launched publicly out of Y Combinator's Summer 2026 batch, founded by Beknazar Abdikamalov — previously co-founder and CTO of the voice-AI company Hupo — with a small San Francisco team.

The benchmark data is the core of the pitch. Speko measures dozens of speech and language models across roughly ten languages — English, Arabic, French, German, Hindi, Norwegian, Spanish, Tamil, and Telugu among them — on word error rate, finalize latency (time from the caller stopping to a finalized transcript), time to first token, and cost per minute, and publishes the results openly at benchmarks.speko.ai. The argument the data makes: English-only leaderboards mislead you, because no single model is best in every language — the top transcription model for Arabic is often not the top one for Spanish. Routing per language, rather than picking one vendor globally, is how you get accuracy everywhere.

Integration is meant to be a base-URL swap. Speko works with the common real-time voice stacks — LiveKit and Pipecat — and exposes an MCP connection for Claude and Cursor, so an existing agent can adopt it without a rewrite. You can also build a new voice agent from scratch on the platform with model selection handled automatically. Speko lists STT pricing that scales with accuracy (roughly a fraction of a cent up to a couple of cents per minute depending on the model's error rate) and cites SOC 2 Type II, HIPAA, and GDPR compliance. Confirm live model coverage, per-model rates, and language support on speko.ai and benchmarks.speko.ai, since the catalog and numbers move.

The honest boundary: Speko is developer infrastructure for building real-time voice applications — phone agents, transcription pipelines, voice assistants. It routes and benchmarks the models; it is not a content tool. It does not write posts in your brand voice, generate persona or avatar video, build carousels, reframe media per platform, or schedule and publish to a social feed. It gets a voice application the best speech model for the moment; turning your product or expertise into distributed content is a separate job.

What you can make with it

  • Real-time voice agents (phone, support, assistant) that run on the best-benchmarked STT, LLM, and TTS models for each language automatically
  • Multilingual transcription pipelines that route each language to its top-performing speech-to-text model instead of one global default
  • Speech-to-speech and text-to-speech that switches vendors by latency, cost, or quality rules without a code change
  • A single billing relationship and one API key across ElevenLabs, OpenAI, Deepgram, AssemblyAI, Cartesia, and other voice providers
  • Open, language-by-language model comparisons (WER, latency, cost) via the public benchmarks at benchmarks.speko.ai
  • Voice apps wired into LiveKit or Pipecat, or connected to Claude and Cursor over MCP, with model selection handled for you

How Kompozy turns Speko output into content

Speko answers a developer's question — "which speech model should this session run on, in this language, right now." It has nothing to say about marketing the thing you built with it. And that is the gap most voice-AI founders and agencies fall into: they ship an impressive multilingual agent and then have no content engine to turn demos, docs, and product updates into an audience. [Kompozy](/) is that engine. Point it at a demo recording, a changelog, or a founder's talk and it produces a week of posts: a [Persona Short](/glossary/persona-shorts) where your face-locked avatar explains what the agent does, a brand-exact [Carousel](/glossary/hyperframes) of the benchmark story, [Clipped Shorts](/glossary/output-buckets) from a longer demo, Quote Graphics from the sharpest lines, plus a launch Blog Article and Email Newsletter — then reframes each to 9:16, 1:1, and 16:9 and publishes across the eight social platforms plus blog and email from one queue.

The neat part is that Kompozy already owns the voice layer you would otherwise wire Speko for: its persona video runs on HeyGen's native text-to-speech, so you get a talking, on-brand avatar without integrating a TTS provider yourself. Speko is where multilingual voice *applications* get the best model per language; Kompozy is where a brand's voice *content* gets generated and shipped, governed by a [Persona Brief](/glossary/persona-brief) and a per-post review gate before it goes out on [Autopilot](/glossary/autopilot). Build the agent on Speko; build the audience on Kompozy.

  1. Build or route your voice agent on Speko so each language runs on its best-benchmarked STT, LLM, and TTS model.
  2. Record a product demo, a walkthrough, or a founder update as your source material.
  3. Drop it into Kompozy as the seed for a batch — a script, a transcript, or the raw clip.
  4. Let Kompozy generate a persona short, a carousel, clipped shorts, quote cards, a blog, and a newsletter in your brand voice.
  5. Reframe each asset per platform and schedule the full set across the nine connected destinations, or hand it to Autopilot behind a review gate.

Frequently asked questions

What is Speko?

Speko is a routing platform for voice AI — the "OpenRouter for voice." One OpenAI-compatible API key sends each session to the speech-to-text, language, and text-to-speech models its benchmarks say win for your language, latency, cost, and quality. It launched out of Y Combinator's Summer 2026 batch.

How does Speko choose which voice model to use?

Speko benchmarks dozens of speech and language models across roughly ten languages on word error rate, finalize latency, time to first token, and cost per minute, and publishes the results at benchmarks.speko.ai. It routes each session to the model that measures best for your language and rules; when the numbers change, the routing changes with no code edit.

Which voice providers does Speko route to?

It routes across many speech-to-text, text-to-speech, and speech-to-speech providers — including ElevenLabs, OpenAI, Deepgram, AssemblyAI, Cartesia, and Google — behind one key. It works with LiveKit and Pipecat and exposes an MCP connection for Claude and Cursor. Confirm the current model list on speko.ai.

Can Speko create or publish social content?

No. Speko is developer infrastructure for building real-time voice applications — it routes and benchmarks speech models and stops there. It has no brand-voice writer, no video or carousel generation, and no scheduler. To turn a voice product or expertise into captioned, on-brand posts across platforms, use a content engine like Kompozy.

How can I use Speko with Kompozy?

Build your voice agent on Speko, then bring a demo, transcript, or update into Kompozy. Kompozy spins it into a persona short, carousel, clipped shorts, quote cards, a blog, and a newsletter under your Persona Brief and publishes them across nine destinations. Kompozy's persona video already ships with native voice, so you can market a voice product without wiring a TTS provider yourself.

Related tools

  • OpenRouterA unified, OpenAI-compatible API that routes one endpoint to 400+ large language models from dozens of providers — with automatic fallback, cost and speed routing, and a single shared credit balance.
  • EchoAn open-weight LLM router from Tracer that routes each request across a pool of open models and, on its evaluated tasks, reaches Claude Fable 5-level quality at roughly a third of the cost — one OpenAI-compatible endpoint for chat, code, and agents.
  • SpeechifyA text-to-speech platform built around low-latency streaming voice — its Simba models turn any script into natural narration for reading, voiceover, and developer apps.
  • Smallest.aiA voice-AI company building ultra-low-latency, human-sounding speech — the Lightning and Waves text-to-speech models, the Pulse speech-to-text stack, and the Atoms real-time voice-agent platform, tuned for sub-100ms conversational voice.
  • OpenAI Voice Models (2026 update)OpenAI's 2026 voice stack in the API — real-time speech-to-speech agents, live translation, streaming transcription, and steerable text-to-speech, all in one family.

← All AI tools · Get started →