Fish Audio builds voice models that can generate speech from text. The company has already released five models in the past year - four for speech generation and one for speech-to-text. Three of those speech models are open source. The newest one, called S2.1 Pro, is only available through a paid API.
The startup's open-source project, Fish Speech, has more than 31,000 stars on GitHub.
The company's first voice generation model was built by former NVIDIA researcher Shijia Liao using a single GPU before being made open source. That initial release sparked a developer community that helped refine the technology and drove adoption. Fish Audio's co-founder Liao brought deep technical expertise to the venture, and the open-source approach allowed the startup to build trust early on by letting users inspect and run the models themselves.
Get the market news that matters in a five-minute read with Market Briefs, our free daily newsletter
By making its models open-source from the start, Fish Audio built a dedicated developer community that keeps refining the technology and expanding its use in gaming, customer service, and content creation.
Today, Fish Audio's voice library offers more than 15,000 natural language controls. CEO Rissa Cao explained the range of use cases this way: "Companies like HeyGen, which use our voices to power AI avatars, want realism in voices; a gaming studio would want expressive voices for their characters; and voice agent companies like LiveKit want more natural-sounding and low-latency voices that are expressive enough for calls."
Fish Audio says it can remove a voice from its platform in less than three minutes after a verified takedown request from the person who owns that voice.
Coreline Ventures partner Osuke Honda said, "A community-centric approach can only become a durable advantage if creators trust the platform. That means consent, transparency, and attribution must be built into the product rather than treated as afterthoughts."
Many competitors - including ElevenLabs, WellSaid, Cartesia, Speechify, Async (formerly Podcastle), and Krisp - are vying for share in the speech generation market. According to 359 Capital partner Rico Mallozzi, "I think what they've been able to build, state-of-the-art models, with the team they have, compared to some of these other well-funded AI labs or companies, is incredible. It shows their technical acumen in closing the gap between artificial-sounding and human-like voices."
The $52 million seed round is exceptionally large for such an early stage, reflecting strong investor belief in Fish Audio's technology and community-driven approach. The company's open-source models have attracted a large developer following, which in turn has driven rapid product iteration and adoption. With the upcoming audio understanding and speech-to-speech models, Fish Audio aims to expand beyond text-to-speech into a broader voice AI platform. Competitors like ElevenLabs and others also have significant funding, but Fish Audio's focus on transparency and consent could be a differentiator in an industry facing scrutiny over voice cloning ethics.
Join Market Briefs, our free daily newsletter, for a quick daily rundown of the markets
