Fish Audio review (2026): honest verdict on the expressive AI voice platform — cloning, S2.1 Pro, open-source Fish Speech, pricing, and who should use it.
Fish Audio is one of the more impressive expressive voice platforms of 2026: word-level emotion control, fast voice cloning, a genuinely open-source lane (Fish Speech), and a hosted flagship (S2.1 Pro), backed by real traction — a $52M seed, ~$21M ARR, and 8M+ users. The caveats are scope and verification: it is a voice engine only — no content generation, no publishing — and some headline metrics are vendor claims worth testing on your own scripts. As a voice platform, it earns a high score.
Fish Audio spent its first year going from an open-source passion project to one of voice AI's faster-growing companies. It began as Fish Speech — an open text-to-speech model started by co-founder and chief scientist Shijia Liao, a former NVIDIA researcher who reportedly trained early versions on a single gaming GPU — and grew, with CEO Rissa Cao, into a company that on July 28, 2026 announced a $52 million seed round and disclosed roughly $21 million in ARR and more than 8 million users. The open-source repository has more than 31,000 GitHub stars.
This review scores Fish Audio as what it is: an AI voice platform. It does expressive text-to-speech, voice cloning, and voice agents; it is not a content-creation suite, and I don't grade it as one — no captioning, no image or video generation, no scheduling. Where it competes, against other voice engines, it competes near the front, and the scores below reflect that.
Two things anchor the verdict. First, the product is genuinely strong on the axes that matter for a voice tool — expressiveness and control (the company advertises more than 15,000 natural-language controls), cloning from a short sample, real-time latency, and a rare combination of an open-source lane plus a hosted flagship, S2.1 Pro. Second, the honest limits: some headline numbers (a claimed listener preference for S2.1 Pro, control counts, model lineup) are the company's own figures without independent benchmarks, so test on your own audio; and the tool's job ends at the voice — everything downstream of the audio file is yours to build.
Everything below reflects Fish Audio's state around its July 28, 2026 funding announcement. Fast-moving startups change models and pricing quickly, so confirm the current lineup and rates on fish.audio before you commit.
Fish Audio is an AI voice platform with two lanes. The open-source lane is the Fish Speech model family, free to download and self-host. The hosted lane is a paid platform and API built around expressiveness: word-level emotion and delivery control, low-latency real-time generation, and voice cloning that can build a usable voice from a short reference clip. The company says it has released several models over the past year — four speech-generation models and one speech-to-text model — open-sourcing three of the speech-generation models while keeping its latest flagship, S2.1 Pro, exclusive to the paid API. It targets a broad audience — indie developers, game designers, creators, and enterprises — and names customers including HeyGen and Sanas, meaning some products you already use may run Fish voices under the hood. It offers a large shared voice library, a free personal tier, and paid subscription and pay-as-you-go plans, with planned additions including an audio-understanding model and a speech-to-speech model. What it does not do is generate written content, images, or video, hold a brand voice across a week of posts, or publish anything — it makes voice, and stops there.
Fish Audio fits three groups cleanly. First, developers building voice features and agents who want low-latency real-time TTS and the option to self-host an open model. Second, creators who need an expressive, clonable voice for a specific job — a faceless-video narration track, an audiobook, a branded read of a script — and want more emotional range than a flat synthetic voice. Third, enterprises that want a hosted voice API with control and cloning behind their own products. Where it fits poorly is anyone expecting a finished content tool: Fish Audio makes audio, not captioned video, carousels, blogs, or published posts, and it has no brand-voice-across-a-week layer, no scheduler, and no publishing. If your bottleneck is producing and distributing finished content rather than synthesizing speech, Fish Audio is one component, not the whole kitchen.
| Dimension | Score | Why |
|---|---|---|
| Voice quality & expressiveness | 4.4 / 5 | Strong, emotive output is the core pitch; the company cites a listener preference for S2.1 Pro over rivals, but that is a vendor blind-test claim — benchmark on your own scripts. |
| Voice cloning | 4.3 / 5 | Cloning from a short reference clip is a headline feature; confirm current sample length and fidelity on the site before relying on it. |
| Control & emotion direction | 4.5 / 5 | Word-level emotion and delivery control (an advertised 15,000+ natural-language controls) is unusually granular for a voice engine. |
| Real-time latency / voice agents | 4.2 / 5 | Low-latency real-time generation suits agents and interactive apps; named customers like HeyGen and Sanas point to production use. |
| Open-source & flexibility | 4.4 / 5 | A rare combination — three open-sourced speech-generation models (Fish Speech, 31,000+ stars) to self-host, plus a hosted flagship for the newest quality. |
| Language coverage | 3.9 / 5 | Broad multilingual support across cloning and TTS; exact language counts vary between the marketing site and press materials, so verify for your target languages. |
| Ease of use & setup | 4.0 / 5 | The hosted API is straightforward; self-hosting Fish Speech is an engineering task, so the effort depends on which lane you pick. |
| Value & pricing | 4.0 / 5 | A free personal tier and open models lower the floor; paid subscription and pay-as-you-go scale with use — reasonable, but confirm current rates. |
| Content-workflow scope | 1.5 / 5 | Voice only — no written content, images, captions, carousels, scheduling, or publishing. Not what the platform is for. |
Fish Audio's pricing spans a wide floor and a scaling ceiling, which is a genuine strength. At the floor, the open-source Fish Speech models are free to self-host, and the hosted platform offers a free personal tier — so you can evaluate and even run low-volume work at no cost. Above that, the paid platform uses subscription and pay-as-you-go plans, with the newest flagship, S2.1 Pro, reserved for the paid API. That structure lets a hobbyist start free, a creator pay for cloning and generation minutes, and an enterprise buy the top model at scale. Because rates on a fast-moving voice platform change, confirm the current subscription and usage pricing on fish.audio before you budget.
The honest asterisk applies to the open-source lane: "free" is a license fact, not a total-cost fact. Self-hosting Fish Speech means supplying the compute and building and maintaining the pipeline that feeds it text and handles the audio — real engineering time. For a developer that is part of the appeal; for a creator who just wants a voice, the hosted API is the sane default and the number to compare is per-minute or per-character usage, not "free vs paid."
The read: as a voice platform, Fish Audio is priced sensibly, with an unusually generous free and open floor and a clear paid path for quality and scale. What that price does not include is any of the content-production work around the voice — writing the copy, making the visuals, assembling the video, or publishing anything. That is not a criticism; it is a scope reminder. You are paying for excellent, controllable voice, and only voice.
| Use case | Fit | Why |
|---|---|---|
| Expressive narration for a faceless video or audiobook | Strong | Word-level emotion control and cloning are exactly what a narration track needs, and Fish Audio is built for it. |
| Building a voice agent or interactive app | Strong | Low-latency real-time TTS and a self-hostable open model make it a strong fit for custom voice integrations. |
| Cloning a consistent house voice from a short sample | Strong | Fast voice cloning is a headline feature; verify current sample length and fidelity for your use. |
| Running a private or high-volume voice at low cost | OK | The open-source Fish Speech lane self-hosts for free, but you supply the compute and the pipeline. |
| Generating written posts, scripts, or blogs | Weak | Fish Audio reads text; it does not write it. There is no copy generation or written brand-voice layer. |
| Producing visual content (carousels, quote cards, video) | Weak | The platform is audio-only; it generates no images or video assets. |
| Scheduling and publishing across platforms | Weak | No scheduler and no social connections — Fish Audio publishes nothing. |
| Running a whole multi-format content operation | Weak | A voice engine covers one ingredient; the copy, visuals, video, and distribution are all still ahead of you. |
To be clear where I stand: I run Kompozy, and Kompozy is not a Fish Audio competitor. Fish Audio is a voice engine — expressive TTS, cloning, voice agents. Kompozy is a content operation you run. I include this note because a fair number of people find a voice tool like Fish Audio while trying to solve a content-volume problem, and it's worth saying plainly that a voice model won't solve that. Great audio is still just audio; you still need something to write the on-brand copy, generate the visuals and video, and get it all published.
That's the honest line between the two. If you want expressive, clonable voice — for narration, an audiobook, or a voice agent — Fish Audio is a genuinely strong pick and this review scores it as one. If your bottleneck is turning one idea into a week of on-brand posts across nine platforms — copy under a Persona Brief, short-form and avatar video (with its own built-in TTS), carousels, quote cards, a blog, and a newsletter, scheduled and published from one queue — that's a content engine's job, and it's the job Kompozy is built for. The clean pairing many creators land on: Kompozy to generate and ship the content, Fish Audio to voice the written scripts into an expressive audio channel. Two tools, two halves — and for social talking-head video, Kompozy's built-in voice means the second tool is optional.
For voice, yes. It is a strong expressive TTS and voice-cloning platform with word-level control, a real open-source lane (Fish Speech), a hosted flagship (S2.1 Pro), and proven traction — a $52M seed, ~$21M ARR, and 8M+ users. It is not worth it as a content-creation tool, because it generates no written posts, images, or video and publishes nothing; it is a voice engine, not a workflow.
Both are strong. Fish Audio's edge is granular emotion control and an open-source lane you can self-host; the company also cites a listener preference for its S2.1 Pro model, though that is a vendor claim. ElevenLabs is a mature, widely benchmarked hosted leader with broad language coverage. Benchmark both on your own scripts — Fish Audio for control and self-hosting, ElevenLabs for a proven managed service.
Partly. The Fish Speech models are open source and free to self-host, and the hosted platform has a free personal tier. Commercial use and the newest flagship model (S2.1 Pro) require paid subscription or pay-as-you-go plans. Confirm current rates on fish.audio, since pricing on a fast-moving voice platform changes.
Yes. Fish Audio supports voice cloning that can build a usable voice from a short reference clip, alongside expressive TTS with word-level emotion controls. Confirm the current required sample length and output fidelity on its site before relying on it for production work.
No. Fish Audio generates voice only — narration, cloned voices, real-time speech. It does not write posts, make images or video, caption clips, or schedule and publish to any platform. For that you need a content engine like Kompozy, which many creators pair with a Fish Audio voice.
Fish Audio was co-founded by chief scientist Shijia Liao, a former NVIDIA researcher, and CEO Rissa Cao. It began as the open-source Fish Speech project and, on its first anniversary (July 28, 2026), announced a $52 million seed round led by Coreline Ventures and Capital Today, disclosing roughly $21 million in ARR and more than 8 million users.
For many creators, yes. Use Kompozy to generate and publish the content — posts, video, blog, newsletter — across platforms, and use Fish Audio to voice the written scripts into an expressive, cloned-voice audio channel like a faceless-video narration or a listen-along blog. They cover two different halves of the job.
Developers building voice features and agents, creators who need an expressive or cloned voice for narration or audiobooks, and enterprises wanting a controllable hosted voice API. It fits poorly for anyone whose real need is producing and distributing finished, on-brand content across platforms — that is a content engine's job.