Fish Audio raises $52M seed to build AI voice models for creators and enterprises

7:00 AM PDT · July 28, 2026

The market for AI-generated voice models is massive. Creative use cases require AI voice models to be more expressive, while enterprises looking to automate customer support and sales ops need them to be more steerable.

Palo Alto-based Fish Audio wants to cater to all of those use cases with its library of more than 15,000 natural language controls. Since launching last year, the startup now has more than 8 million people using the open-source or hosted versions of its models, and generates annual recurring revenue of $21 million.

To continue building on that traction, the startup on Tuesday said it has raised $52 million in a seed round that was led by Coreline Ventures and Capital Today. The funding also saw participation from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.

(Note: TechCrunch updated the story to reflect that the company raised 50M.)

Fish Audio started as a small project by former NVIDIA researcher Shijia Liao, who, frustrated by non-expressive synthetic voices available on the market, trained a voice generation model on a single GPU, which he then open-sourced. The Fish Speech repository on GitHub now has more than 31,000 stars and is used by indie developers, video game designers, and creators.

The company has launched five models in the last year: four speech generation models and one speech-to-text model. It has open-sourced three of its speech generation models, but its latest S2.1 Pro model is available only through its paid API.

Fish Audio offers paid monthly plans suited for creators and teams that unlock a set number of minutes of generation plus voice cloning features. The company also offers an enterprise version of its APIs and platform, and says organizations like HeyGen and Sanas are already using it.

CEO and co-founder Rissa Cao said different enterprise customers want different traits: HeyGen wants realism for AI avatars; gaming studios want expressive character voices; LiveKit-style voice agents want natural, low-latency voices.

One way the startup has built its library of voices is by asking users to submit their own voices for training its models, and compensating them if their voices are used. That resulted in some trouble a few months ago, as some creators alleged that their voices were uploaded to Fish Audio without their consent. The startup had a DMCA takedown process; Cao said takedowns are now automated to under three minutes with a voice sample or contract as proof. Consent/ownership issues remain industry-wide.

Osuke Honda (Coreline Ventures) stressed that community-driven models only work with creator trust: consent, transparency, attribution, verified voice ownership, clear licensing, and eventually revenue-sharing.

Looking ahead, Fish Audio plans to release an audio understanding model this year and is building a speech-to-speech model. Competitors named include ElevenLabs, WellSaid, Cartesia, Speechify, Async (previously Podcastle), and Krisp.