Learn About Fish Audio's AI Products
Create a free account to access in-depth lessons on each tool and model.
Start Learning Free📋About Fish Audio
Updated July 29, 2026Fish Audio is a speech AI company building text-to-speech, voice-cloning and speech-to-text models for creators and enterprises. It was founded in 2025 by Shijia Liao, a former NVIDIA researcher, and chief executive Rissa Cao, and it closed a $52 million seed round in July 2026, led by Coreline Ventures and Capital Today, with participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners and HF0.
The company ships on two tracks at once. Its Fish Speech research models are published on GitHub, where the repository has passed 31,000 stars, and are built on a dual-autoregressive architecture trained on more than 10 million hours of audio across 80-plus languages. Its commercial platform layers hosted products on top of that research: real-time text-to-speech built on the S2.1 Pro model, 15-second voice cloning, speech-to-text with multispeaker and emotion tagging, a voice changer, audio separation and translation, and a Story Studio aimed at audiobook production. Fish Audio reports more than 8 million users of its open and hosted versions, and its voice library exceeds 2 million community voices.
The published models are widely described as open source, but that framing needs a caveat. Fish Speech ships under the Fish Audio Research License, which permits research, hobbyist and evaluation use royalty-free while requiring a separate written agreement from Fish Audio for any commercial purpose — including hosted applications, APIs and internal business use. The license also requires "Built with Fish Audio" attribution and forbids using model outputs to train competing foundation models. Teams that want commercial rights without negotiating a bilateral agreement generally use the paid hosted tiers, which do grant commercial use, or choose a permissively licensed alternative.
🛠️Products & Tools (1)
Speech platform with real-time text-to-speech, 15-second voice cloning, speech-to-text and audio tooling, backed by the Fish Speech research models. Free tier plus paid plans from $5.50 per month; the downloadable weights are research-licensed, not permissively open.
