Learning Objectives
- Understand what Fish Audio offers and how its hosted platform relates to its downloadable models
- Evaluate the Fish Audio Research License and what it does and does not permit
- Compare Fish Audio against ElevenLabs, OpenAI TTS and genuinely open alternatives
- Decide which tier fits a creator, a small team, or an enterprise deployment
What Is Fish Audio?
Fish Audio is a speech AI platform from a company of the same name, founded in 2025 by former NVIDIA researcher Shijia Liao and chief executive Rissa Cao. It raised a $52 million seed round in July 2026, led by Coreline Ventures and Capital Today.
The product covers most of the speech stack in one place: real-time text-to-speech built on the S2.1 Pro model, voice cloning from a 15-second sample, speech-to-text with multispeaker and emotion tagging, plus a voice changer, audio separation, audio translation, sound effects, and a Story Studio aimed at audiobook production. It supports 30-plus languages, and its community voice library has passed 2 million voices.
What distinguishes it from most competitors is the second track: the underlying Fish Speech models are downloadable from GitHub, where the repository has passed 31,000 stars. They use a dual-autoregressive architecture trained on more than 10 million hours of audio across 80-plus languages, with the flagship S2-Pro at roughly 4 billion parameters.
⚠️Warning
"Open source" here does not mean what it usually means. Fish Speech ships under the Fish Audio Research License, not MIT or Apache 2.0. Research, hobbyist and evaluation use is royalty-free, but any commercial purpose — including hosted apps, APIs and internal business use — requires a separate written agreement from Fish Audio. The license also mandates "Built with Fish Audio" attribution and forbids using outputs to train competing foundation models. If you need commercial rights without negotiating a contract, use the paid hosted tiers (which do grant commercial use) or pick a permissively licensed model.
Key Capabilities
Voice Cloning
Fish Audio clones a voice from roughly 15 seconds of reference audio. Free accounts get 3 public voice slots; paid tiers add private slots and a smaller number of "professional" slots for higher-fidelity clones.
Real-Time Text-to-Speech
The S2.1 Pro model is positioned around expressiveness and emotional control rather than raw intelligibility, and it streams in real time — which is the requirement that matters for voice agents and interactive applications, as opposed to batch narration.
Speech-to-Text with Emotion Tags
The transcription model handles multiple speakers and emits emotion tags alongside the text. That is a genuinely useful combination for podcast and interview workflows, where knowing who spoke and how is often as valuable as the transcript.
💡Key Concept
Credits versus minutes. Fish Audio meters in credits, and each plan also advertises an approximate generation-minutes ceiling. Treat the minutes figure as the practical number — credits convert at different rates across models, so the minutes estimate is the one to budget against.
Pricing
- 8,000 credits (about 7 minutes)
- 3 public voice slots
- Commercial use allowed
- 250,000 credits (about 200 minutes)
- 10 private voice slots
- 1 professional voice slot
- 2,000,000 credits (about 1,620 minutes)
- 3 team seats
- 5 professional voice slots
- 25,000,000 credits (about 6,250 minutes)
- 10 team seats
- 15 professional voice slots
- Custom volume
- Zero data retention; on-premise
- SOC 2 available
Annual billing takes about 33 percent off every paid tier. Note that commercial use is granted on the free tier — unusual in this category, and a meaningful difference from the research license attached to the downloadable weights.
Fish Audio vs. Other Speech Platforms
| Tool | Provider | Weights available | License for commercial use | Key strength |
|---|---|---|---|---|
| Fish Audio | Fish Audio | Yes (research license) | Separate agreement required | Full speech stack in one platform; cheap paid tiers |
| ElevenLabs | ElevenLabs | No | Hosted only | Highest expressiveness; deepest voice library |
| OpenAI TTS / Whisper | OpenAI | Whisper yes, TTS no | Whisper is MIT | Broad language coverage; Whisper is genuinely open |
| Voxtral TTS | Mistral AI | Yes | Open license | Runs on consumer hardware; European AI |
| Murf AI | Murf | No | Hosted only | Studio workflow for marketing voiceover |
Strengths
- Breadth in one platform — text-to-speech, cloning, transcription, translation and separation without stitching several vendors together
- Aggressive pricing — the Plus tier at five dollars fifty per month undercuts most hosted competitors substantially
- Commercial use on the free tier — rare in this category, and it lowers the cost of a real evaluation
- Fast voice cloning — a 15-second reference sample is at the short end of what the category requires
- Real weights to inspect — even under a restrictive license, being able to read and run the model is more than most hosted competitors allow
- Multilingual reach — 30-plus languages on the platform, with the research models trained across 80-plus
Limitations and Considerations
- The license is the headline caveat — "open source" in most coverage is misleading; commercial self-hosting requires a bilateral agreement with Fish Audio
- Young company — founded in 2025, so track record, support and roadmap stability are unproven relative to incumbents
- Attribution requirement — research-license users must display "Built with Fish Audio," which not every project can accommodate
- Credit accounting adds friction — budgeting requires translating credits into minutes per model
- Voice-cloning ethics and consent — 15-second cloning is powerful and easy to misuse; confirm you have permission for any voice you replicate
- Enterprise features are gated — zero data retention, on-premise deployment and SOC 2 sit behind custom annual pricing
Company Details
| Detail | Info |
|---|---|
| Developer | Fish Audio (San Francisco, California) |
| Founded | 2025 |
| Founders | Shijia Liao (ex-NVIDIA) and Rissa Cao (CEO) |
| Funding | $52 million seed, July 2026, led by Coreline Ventures and Capital Today |
| Flagship model | S2.1 Pro (hosted); S2-Pro at about 4 billion parameters (downloadable) |
| Weights license | Fish Audio Research License (non-commercial without separate agreement) |
| Languages | 30-plus on the platform; 80-plus in research training data |
| Website | fish.audio |
Key Takeaways
- Fish Audio covers the whole speech stack — text-to-speech, 15-second voice cloning, transcription, translation and separation — on one platform
- Its Fish Speech weights are downloadable, but under a research license that requires a separate written agreement for any commercial use
- Paid plans start at five dollars fifty per month and commercial use is granted even on the free tier, which makes evaluation unusually cheap
- It raised a $52 million seed round in July 2026 and reports more than 8 million users across its open and hosted versions
- Choose it for breadth and price; choose ElevenLabs for maximum expressiveness, or a permissively licensed model if you must self-host commercially


