// AI TOOLS · FISH AUDIO

Fish Audio

An AI voice platform for expressive real-time text-to-speech and fast voice cloning — with an open-source model family (Fish Speech) and a hosted flagship, S2.1 Pro, aimed at creators, developers, and enterprises.

KompozyTurn one idea into a week of content — across every platform, published for you.
Get Started →

Last verified · 2026-07-28 · by Moe Ameen

What Fish Audio is

Fish Audio is an AI voice company that builds text-to-speech, voice cloning, and voice-agent tools. It started as an open-source passion project — Fish Speech, whose GitHub repository has more than 31,000 stars — begun by co-founder and chief scientist Shijia Liao, a former NVIDIA researcher, and grew into a company co-founded with CEO Rissa Cao. On July 28, 2026, roughly a year after launch, Fish Audio announced a $52 million seed round led by Coreline Ventures and Capital Today, and disclosed that it had reached about $21 million in annual recurring revenue with more than 8 million users.

The product runs in two lanes. The open-source lane is the Fish Speech model family, which you can download and self-host for free. The hosted lane is a paid platform and API: the company says it has released several voice models over the past year — four speech-generation models and one speech-to-text model — open-sourcing three of the speech-generation models while keeping its latest flagship, S2.1 Pro, exclusive to the paid API. The hosting pitch is expressiveness and control: word-level emotion and delivery direction (the company advertises more than 15,000 natural-language controls), low-latency real-time generation, and voice cloning that can build a usable voice from a short reference clip.

Fish Audio positions itself for a broad audience — indie developers, game designers, creators, and enterprises — and names companies like HeyGen and Sanas among its customers, meaning some products you already use may run Fish voices under the hood. It also offers a large shared voice library and a free tier for personal use, with paid subscription and pay-as-you-go options for commercial work. Planned additions include an audio-understanding model and a speech-to-speech model.

The honest framing: Fish Audio is a voice engine, not a content pipeline. It turns text into expressive speech and can clone a voice — but it writes no script, edits no video, designs no post, holds no brand voice across a week of content, and publishes nothing. What you get back is an audio file (or a streaming voice), and everything around it is yours to build.

What you can make with it

  • Expressive text-to-speech narration with word-level emotion and delivery control
  • A cloned brand or persona voice built from a short reference sample
  • Real-time, low-latency voice for agents, assistants, and interactive apps
  • Multilingual voiceover for faceless videos, explainers, and audiobooks
  • Speech-to-text transcripts, plus voices pulled from a large shared library
  • A self-hosted, open-source voice via the Fish Speech models for private or high-volume use

How Kompozy turns Fish Audio output into content

The most interesting thing Fish Audio gives a brand is a clonable, expressive voice — build one usable house voice from a short sample, then generate narration that actually carries emotion. The gap is that a cloned voice on its own is just an audio track: it has no script worth reading, no video to sit inside, and no way onto a feed. Kompozy supplies exactly those missing pieces. Kompozy writes the words Fish Audio speaks — the video script, the Text Posts, the Blog Article, the Email Newsletter — all governed by your Persona Brief and banned-word filters, so the narration you clone in Fish is reading on-brand copy rather than something you had to hand-write. You voice that script in Fish Audio, then drop the audio into a faceless Listicle or Naturalistic Video, a slideshow, or a listen-along blog, and let Kompozy fan the same idea into the visual and text formats and publish the lot across nine platforms.

There's a neat pairing on identity. Fish Audio locks a consistent *voice*; Kompozy locks a consistent *face* — Gemini face-locked Persona Photos and Persona Tweets, plus HeyGen-driven Persona Shorts and Persona Frames — so together you get one recurring persona whose look and sound both stay on-brand across every post. And because Kompozy's avatar video already ships with HeyGen's native TTS (and HeyGen is itself a Fish Audio customer), you can keep it simple for talking-head social and reserve your Fish clone for the standalone narration, audiobook, and voice-agent work where its expressiveness earns its place. Fish makes the voice; Kompozy writes what it says and turns it into a published, cross-platform content operation.

  1. In Fish Audio, clone your house voice from a short reference clip (or pick one from its voice library).
  2. In Kompozy, generate the source — a Blog Article and a tight narration script — held to one voice by your Persona Brief.
  3. Run that script through Fish Audio to produce an expressive voiceover, tuning emotion and pacing with its controls.
  4. Assemble the audio into a faceless Listicle/Naturalistic Video or a listen-along blog audio version.
  5. Back in Kompozy, fan the same idea into Clipped and Persona Shorts, Carousels, Quote Graphics, and a Newsletter, then schedule and publish across nine social platforms plus blog and email with Autopilot.

Frequently asked questions

What is Fish Audio?

Fish Audio is an AI voice platform for expressive text-to-speech, voice cloning, and voice agents. It grew from the open-source Fish Speech project (31,000+ GitHub stars) into a company that offers self-hosted open models plus a paid hosted API, including its flagship S2.1 Pro model. On July 28, 2026 it announced a $52 million seed round and disclosed about $21 million ARR and 8 million-plus users.

Is Fish Audio open source or paid?

Both. The Fish Speech model family is open source and free to self-host, and the company says it has open-sourced three of its speech-generation models. Its newest flagship, S2.1 Pro, is available only through the paid hosted API, which also offers a free personal tier and paid subscription and pay-as-you-go plans for commercial use.

Can Fish Audio clone a voice?

Yes. Fish Audio supports voice cloning that can build a usable voice from a short reference clip, alongside expressive text-to-speech with word-level emotion and delivery controls (the company advertises more than 15,000 natural-language controls). Confirm current cloning length and quality on its site before relying on it for production.

Does Fish Audio create or publish social content?

No. Fish Audio generates voice — narration, cloned voices, real-time speech — and stops there. It writes no script, makes no video or image, captions nothing, and publishes to no platform. For that you pair it with a content engine like Kompozy, which generates the copy, visuals, and video and publishes across nine platforms.

How do I use Fish Audio with Kompozy?

Use Kompozy to write the on-brand script and long-form content under your Persona Brief and to publish everything across platforms, then run the script through Fish Audio to voice an expressive, cloned-voice narration track for a faceless video or a listen-along blog. Kompozy owns generation and distribution; Fish Audio owns the voice layer.

Related tools

  • HeyGenAI avatar video platform that turns a text script into a talking-head video — in 175+ languages.
  • Kokoro TTSAn open-weight, 82-million-parameter text-to-speech model that runs high-quality narration locally on a CPU — free, offline, and Apache-2.0 licensed for commercial use.
  • SpeechifyA text-to-speech platform built around low-latency streaming voice — its Simba models turn any script into natural narration for reading, voiceover, and developer apps.
  • Speechify Simba 3.2 APISpeechify's streaming-native text-to-speech model, exposed as a developer API — sub-300ms latency, prosody-level emotion, and the top spot on the Artificial Analysis TTS Arena.
  • MetaVoiceA production duplex speech model for revenue phone calls — a single AI that listens and speaks at the same time, so conversations survive interruptions, overlap, and background voices instead of taking rigid turns.
  • Grok VoicesxAI's upgraded voice generation for Grok — 21 new flagship voices (26 total), each natively multilingual across 25+ languages and cast for a specific job like support, characters, commentary, advertising, or education.

← All AI tools · Get started →