An AI voice platform for expressive real-time text-to-speech and fast voice cloning — with an open-source model family (Fish Speech) and a hosted flagship, S2.1 Pro, aimed at creators, developers, and enterprises.
Last verified · 2026-07-28 · by Moe Ameen
Fish Audio is an AI voice company that builds text-to-speech, voice cloning, and voice-agent tools. It started as an open-source passion project — Fish Speech, whose GitHub repository has more than 31,000 stars — begun by co-founder and chief scientist Shijia Liao, a former NVIDIA researcher, and grew into a company co-founded with CEO Rissa Cao. On July 28, 2026, roughly a year after launch, Fish Audio announced a $52 million seed round led by Coreline Ventures and Capital Today, and disclosed that it had reached about $21 million in annual recurring revenue with more than 8 million users.
The product runs in two lanes. The open-source lane is the Fish Speech model family, which you can download and self-host for free. The hosted lane is a paid platform and API: the company says it has released several voice models over the past year — four speech-generation models and one speech-to-text model — open-sourcing three of the speech-generation models while keeping its latest flagship, S2.1 Pro, exclusive to the paid API. The hosting pitch is expressiveness and control: word-level emotion and delivery direction (the company advertises more than 15,000 natural-language controls), low-latency real-time generation, and voice cloning that can build a usable voice from a short reference clip.
Fish Audio positions itself for a broad audience — indie developers, game designers, creators, and enterprises — and names companies like HeyGen and Sanas among its customers, meaning some products you already use may run Fish voices under the hood. It also offers a large shared voice library and a free tier for personal use, with paid subscription and pay-as-you-go options for commercial work. Planned additions include an audio-understanding model and a speech-to-speech model.
The honest framing: Fish Audio is a voice engine, not a content pipeline. It turns text into expressive speech and can clone a voice — but it writes no script, edits no video, designs no post, holds no brand voice across a week of content, and publishes nothing. What you get back is an audio file (or a streaming voice), and everything around it is yours to build.
The most interesting thing Fish Audio gives a brand is a clonable, expressive voice — build one usable house voice from a short sample, then generate narration that actually carries emotion. The gap is that a cloned voice on its own is just an audio track: it has no script worth reading, no video to sit inside, and no way onto a feed. Kompozy supplies exactly those missing pieces. Kompozy writes the words Fish Audio speaks — the video script, the Text Posts, the Blog Article, the Email Newsletter — all governed by your Persona Brief and banned-word filters, so the narration you clone in Fish is reading on-brand copy rather than something you had to hand-write. You voice that script in Fish Audio, then drop the audio into a faceless Listicle or Naturalistic Video, a slideshow, or a listen-along blog, and let Kompozy fan the same idea into the visual and text formats and publish the lot across nine platforms.
There's a neat pairing on identity. Fish Audio locks a consistent *voice*; Kompozy locks a consistent *face* — Gemini face-locked Persona Photos and Persona Tweets, plus HeyGen-driven Persona Shorts and Persona Frames — so together you get one recurring persona whose look and sound both stay on-brand across every post. And because Kompozy's avatar video already ships with HeyGen's native TTS (and HeyGen is itself a Fish Audio customer), you can keep it simple for talking-head social and reserve your Fish clone for the standalone narration, audiobook, and voice-agent work where its expressiveness earns its place. Fish makes the voice; Kompozy writes what it says and turns it into a published, cross-platform content operation.
Fish Audio is an AI voice platform for expressive text-to-speech, voice cloning, and voice agents. It grew from the open-source Fish Speech project (31,000+ GitHub stars) into a company that offers self-hosted open models plus a paid hosted API, including its flagship S2.1 Pro model. On July 28, 2026 it announced a $52 million seed round and disclosed about $21 million ARR and 8 million-plus users.
Both. The Fish Speech model family is open source and free to self-host, and the company says it has open-sourced three of its speech-generation models. Its newest flagship, S2.1 Pro, is available only through the paid hosted API, which also offers a free personal tier and paid subscription and pay-as-you-go plans for commercial use.
Yes. Fish Audio supports voice cloning that can build a usable voice from a short reference clip, alongside expressive text-to-speech with word-level emotion and delivery controls (the company advertises more than 15,000 natural-language controls). Confirm current cloning length and quality on its site before relying on it for production.
No. Fish Audio generates voice — narration, cloned voices, real-time speech — and stops there. It writes no script, makes no video or image, captions nothing, and publishes to no platform. For that you pair it with a content engine like Kompozy, which generates the copy, visuals, and video and publishes across nine platforms.
Use Kompozy to write the on-brand script and long-form content under your Persona Brief and to publish everything across platforms, then run the script through Fish Audio to voice an expressive, cloned-voice narration track for a faceless video or a listen-along blog. Kompozy owns generation and distribution; Fish Audio owns the voice layer.