Fish Audio just landed $52 million in seed funding to build AI voice models for both creators and enterprises, the Palo Alto startup announced Tuesday. According to TechCrunch AI, the round was led by Coreline Ventures and Capital Today, with a long list of firms joining in: 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. That’s a large seed by any measure, and it signals how hot the synthetic voice market has become.
What stands out here is the traction behind the check. Fish Audio launched just last year and already claims more than 8 million users across its open source and hosted models, plus $21 million in annual recurring revenue. Most seed-stage startups don’t have those numbers. This one does.
🎙️ From one GPU to 31,000 stars
The company started as a side project. Former Nvidia researcher Shijia Liao got frustrated with the flat, robotic voices on the market, trained a voice model on a single GPU, and open-sourced it. That project, Fish Speech, now has over 31,000 stars on GitHub and gets used by indie developers, game designers, and creators.
Since then, Fish Audio has shipped five models in a year: four for speech generation and one for speech-to-text. Three of the speech models are open source. The newest, S2.1 Pro, is locked behind a paid API. The pitch is control. Fish Audio offers more than 15,000 natural language controls, which lets customers dial in exactly the kind of voice they want.
That flexibility is the whole strategy. As CEO and co-founder Rissa Cao told TechCrunch AI, different customers want different things:
- HeyGen wants realism to power AI avatars
- Gaming studios want expressive voices for characters
- Voice agent companies like LiveKit want natural, low-latency voices built for calls
⚠️ The consent problem
Here’s the part worth watching. Fish Audio builds its voice library partly by asking users to submit their own voices and paying them if those voices get used. A few months ago that backfired. Some creators alleged their voices were uploaded without consent, and the company’s takedown process was slow.
Cao says Fish Audio has now automated it. Creators can submit a voice sample or a contract to prove ownership, and the voice comes off the platform in under three minutes. That’s a real improvement. But it doesn’t stop someone from uploading an artist’s voice in the first place, and the voice stays live until the artist notices and files for removal. The burden still sits on the person being copied.
Coreline Ventures partner Osuke Honda was blunt about what has to change. “A community-centric approach can only become a durable advantage if creators trust the platform,” he said, arguing that consent, transparency, and attribution need to be “built into the product rather than treated as afterthoughts.” He wants the industry moving toward verified voice ownership, clear licensing, easy takedowns, and revenue-sharing when voices get used commercially.
Why this matters
The voice AI market is crowded and getting more so. Fish Audio is stepping into a ring that already includes ElevenLabs, WellSaid, Cartesia, Speechify, Async (formerly Podcastle), and Krisp. Everyone’s chasing the same creator and enterprise budgets.
Fish Audio’s edge, according to 359 Capital partner Rico Mallozzi, is fine-grained developer controls and cost-efficient training. He credits the team with building state-of-the-art models on a fraction of what better-funded labs spend. “It shows their technical acumen in closing the gap between artificial-sounding and human-like voices,” he told TechCrunch AI.
Cao is candid that the company didn’t strictly need the money. Its open source plus creator plans were already running efficiently. What pushed it to raise was ambition: better models, enterprise features, and rising investor interest it decided to meet.
What comes next
Fish Audio plans to release an audio understanding model this year and is building a speech-to-speech model on top of that. If you’re a developer or a business shopping for voice AI, expect a more capable, more steerable option entering the field, and expect the consent question to keep following the whole category.
For the full breakdown, the original report is available at TechCrunch AI.