// AI TOOLS · METAVOICE

MetaVoice

A production duplex speech model for revenue phone calls — a single AI that listens and speaks at the same time, so conversations survive interruptions, overlap, and background voices instead of taking rigid turns.

KompozyTurn one idea into a week of content — across every platform, published for you.
Get Started →

Last verified · 2026-07-24 · by Moe Ameen

What MetaVoice is

MetaVoice is a duplex speech-to-speech model built for real-time phone conversations — sales and support calls where a rigid, turn-based agent loses the caller. The company positions it as "the duplex speech model for revenue calls," and the core idea is that a single model takes in audio and produces speech at the same time, rather than running a separate speech-to-text step, a language model, and a text-to-speech step in sequence. Because it listens while it speaks, the model can handle interruptions, overlapping talk, and background voices without the conversation breaking, which is where cascaded voice-agent stacks tend to stumble. MetaVoice frames the problem it's solving bluntly: a large share of people abandon voice agents within the first 30 seconds of a call.

Two things distinguish the approach. First, the model "reasons over speech, not text" — it works from the actual audio, so it picks up tone and ambient context and can ask a clarifying question when it isn't sure, instead of committing to a bad transcription. Second, there is no pipeline to assemble: no ASR, TTS, turn detector, or dialogue harness to wire together. You define the workflow, the tools it can call, and its personality through prompts. MetaVoice's landing page cites a response time around 350 milliseconds, and the product ships production controls a voice deployment actually needs — guardrails that can filter or block a response before the caller hears it, the ability to inspect the model's reasoning, tool calls, and speech in both text and audio, and observable streams that feed existing monitoring tools. It can be deployed inside a customer's own VPC so call data stays on-premise, and it improves over time by fine-tuning on a customer's real calls.

MetaVoice is the team behind MetaVoice-1B, the ~1.2-billion-parameter open-source text-to-speech model released in early 2024, and the duplex product is a distinct, later direction focused on live two-way conversation rather than one-way narration. A key piece of the work is a proprietary speech-separation model the company built to turn messy, mixed real-world call recordings into the clean, per-speaker tracks a duplex model needs to train on. The honest framing for a creator: MetaVoice is B2B voice-agent infrastructure for phone calls, not a content-generation tool. It produces a live conversation between an agent and a caller — not a caption, a clip, or a post you publish. Treat any specific latency figures, tier details, or dates here as a snapshot and verify current specifics with MetaVoice directly.

What you can make with it

  • A real-time AI phone agent for inbound or outbound calls that holds a natural, two-way conversation instead of a scripted, turn-based exchange
  • Sales and support voice agents that survive interruptions, overlapping speech, and background noise on a live call
  • A voice agent that reasons over what it hears — tone and context — and asks a clarifying question when a caller is ambiguous
  • A tool-using agent whose workflow, callable tools, and personality are defined through prompts rather than a hand-built ASR/LLM/TTS stack
  • A privacy-controlled deployment inside your own VPC, with guardrails and observability, that fine-tunes on your real call data over time

How Kompozy turns MetaVoice output into content

MetaVoice and Kompozy never touch the same output, and that's exactly why they sit well together for a founder or agency operating in the voice-AI space. MetaVoice's product is the call itself — a live conversation with one caller at a time, deployed inside a VPC. Nothing about that is publishable content; it's infrastructure that runs quietly behind a phone number. The audience-facing story of that product — what duplex speech is, why listening-while-speaking beats turn-taking, how it cuts the 30-second hang-up problem — is a separate job, and it's the one Kompozy is built for.

Concretely: if you're launching or reselling a MetaVoice-style voice agent, drop your positioning, a demo transcript, or a founder talk into Kompozy and it fans that single source into a week of on-brand content — Persona Shorts and HeyGen avatar video explaining the duplex-vs-cascaded difference, brand-exact Carousels via HyperFrames breaking down the architecture, Quote Graphics pulling the "40% hang up in 30 seconds" stat, a Blog Article on when duplex models make sense, an Email Newsletter for your pilot list, and native Text Posts — all governed by a Persona Brief so the technical story still reads in your voice. Then Autopilot and a per-post review pipeline schedule and publish the batch across the eight social platforms plus blog and email. MetaVoice answers the phone; Kompozy builds the audience that makes the phone ring.

  1. Bring one source into Kompozy — your product positioning, a recorded demo call transcript, or a founder explainer video on why duplex speech beats turn-based agents.
  2. Set a Persona Brief so every generated piece holds your technical voice and banned words, whether the reader is a prospect or an investor.
  3. Generate the batch: Persona/HeyGen avatar explainers, brand-exact Carousels on the architecture, Quote Graphics for the standout stats, a Blog Article, a Newsletter, and Text Posts — from that one source.
  4. Let each asset reframe to 9:16, 1:1, and 16:9 so one explainer fans into platform-native posts.
  5. Schedule and publish across the eight social platforms plus blog and email from one queue with Autopilot and a per-post review pipeline.

Frequently asked questions

What is MetaVoice?

MetaVoice is a duplex speech-to-speech model built for real-time phone conversations — sales and support calls. A single model listens and speaks at the same time, so it handles interruptions, overlapping talk, and background voices instead of taking rigid turns. It comes from the team behind the open-source MetaVoice-1B text-to-speech model.

What does "duplex" mean for a voice model?

Duplex means the model takes in audio and produces speech simultaneously, rather than strictly alternating turns like a walkie-talkie. That lets it back-channel, absorb interruptions, and keep talking through overlap and background noise — closer to how a real phone call flows than a turn-based ASR-then-LLM-then-TTS pipeline.

Is MetaVoice a content-creation tool?

No. MetaVoice produces live phone conversations between an AI agent and a caller — not captions, clips, carousels, or posts. It is voice-agent infrastructure. To turn the story of a voice-AI product into publishable content across platforms, you would use a generation and publishing engine like Kompozy alongside it.

How is MetaVoice different from MetaVoice-1B?

MetaVoice-1B was the team's ~1.2B-parameter open-source text-to-speech model (one-way narration) from early 2024. The duplex speech model is a later, distinct direction focused on live two-way conversation for phone calls — listening and speaking at once — rather than reading a script aloud.

Can I use MetaVoice and Kompozy together?

Yes, and they don't overlap. MetaVoice runs the live calls; Kompozy generates and publishes the content that markets the product or drives the calls — persona video, carousels, blogs, newsletters, and posts across the eight social platforms plus blog and email, all held to one Persona Brief.

Related tools

  • GPT-LiveOpenAI's new full-duplex voice models for ChatGPT — they listen and speak at the same time, and can hand a question off for a web search mid-conversation.
  • OpenAI Voice Models (2026 update)OpenAI's 2026 voice stack in the API — real-time speech-to-speech agents, live translation, streaming transcription, and steerable text-to-speech, all in one family.
  • SpeechifyA text-to-speech platform built around low-latency streaming voice — its Simba models turn any script into natural narration for reading, voiceover, and developer apps.

← All AI tools · Get started →