The voice-AI startup that began as an open-source passion project — Fish Speech, 31,000-plus GitHub stars — disclosed a $52 million seed round, about $21 million in annual recurring revenue, and more than 8 million users on its first anniversary.
2026-07-28 · by Moe Ameen
On July 28, 2026, Fish Audio announced a $52 million seed round led by Coreline Ventures and Capital Today, with participation from a group of funds including 359 Capital, Play Time, HF0, 645 Ventures, Parable, Carya Venture Partners, and Alphalist Partners. The company timed the disclosure to its first anniversary and said it had grown from zero to roughly $21 million in annual recurring revenue and more than 8 million users across creators, developers, and enterprises.
Fish Audio builds AI voice tools: expressive text-to-speech, voice cloning, and voice agents. It began as an open-source project, Fish Speech, started by co-founder and chief scientist Shijia Liao — a former NVIDIA researcher who, per the company, trained early models on a single gaming GPU — and grew into a company co-founded with CEO Rissa Cao. The open-source repository has more than 31,000 GitHub stars. The company says it has shipped several models over the past year — four speech-generation models and one speech-to-text model — open-sourcing three of the speech-generation models while keeping its latest flagship, S2.1 Pro, exclusive to its paid API.
The hosted pitch is control and expressiveness: word-level emotion and delivery direction (the company advertises more than 15,000 natural-language controls), low-latency real-time generation, and voice cloning from a short reference clip. Fish Audio names customers including HeyGen and Sanas, and says its planned releases include an audio-understanding model and a speech-to-speech model. As always with a fast-moving launch, treat model names, control counts, and headline metrics as a snapshot and confirm current details on the company's own site.
A well-funded voice model doesn't change the creator's real problem: an expressive, perfectly cloned voice is still just an audio track until something writes what it says, wraps it in a video or a post, and gets it onto every feed. As voice commoditizes, the moat moves downstream to exactly that — and that is the layer Kompozy runs. Use Fish Audio to clone your house voice and voice a narration, then let Kompozy do the rest: it writes the script and long-form copy under your Persona Brief and banned-word filters (so the voice is reading on-brand words, not something you hand-wrote), drops the audio into a faceless Listicle or Naturalistic Video, and fans the same idea into Persona Shorts, Carousels, Quote Graphics, a Blog Article, and a Newsletter — then schedules and publishes the whole set across nine social platforms plus blog and email from one queue with Autopilot. Fish locks a consistent voice; Kompozy's face-locked persona locks a consistent look, so together you get one recurring identity that sounds and looks like your brand everywhere.
There's a same-day publishing play, too. "An AI voice startup just raised $52M" is a story your audience is searching right now. Drop your take into Kompozy as a source and it turns one point of view into a captioned short, a carousel on what expressive TTS changes for creators, a blog explainer, a newsletter, and platform-native posts — generated in your voice, scheduled, and published across every channel in an afternoon. Being early and specific on a funding story like this is how a single opinion becomes a full content cycle.
Fish Audio announced a $52 million seed round on July 28, 2026, led by Coreline Ventures and Capital Today, with participation from funds including 359 Capital, Play Time, HF0, and 645 Ventures. It disclosed the raise alongside about $21 million in annual recurring revenue and more than 8 million users, on the company's first anniversary.
Fish Audio builds AI voice tools — expressive text-to-speech, voice cloning from a short reference clip, and voice agents. It offers open-source models (the Fish Speech family) you can self-host plus a paid hosted API, including its flagship S2.1 Pro. It advertises word-level emotion control and low-latency real-time generation, and names customers including HeyGen and Sanas.
It signals that expressive, clonable AI voice is getting cheaper and more available, so a good voiceover stops being a bottleneck — and the scarce work moves downstream to writing the script, making the visuals and video, and publishing everywhere. A cloned voice is a component, not a content operation; a content engine like Kompozy writes the on-brand copy, generates the formats, and publishes across nine platforms plus blog and email.