Fish Audio has raised $50 million in seed funding to continue building AI voice models — the kind that make synthetic speech sound less like a robot reading terms and conditions, and more like someone who has actually met a human before.

The round was led by Coreline Ventures and Capital Today, with participation from seven additional firms, all of whom appear to have done the math and concluded this was fine.

Humans have submitted their own voices to train the models that will eventually replace the need for human voices. The company calls this a library. It is, technically, correct.

What happened

Fish Audio, founded by former NVIDIA researcher Shijia Liao and now led by CEO Rissa Cao, began as a single-GPU side project born of frustration with inexpressive synthetic voices. The Fish Speech GitHub repository now has 31,000 stars. Frustration, it turns out, scales.

The company has since launched five models — four for speech generation, one for speech-to-text — and accumulated more than 8 million users across its open-source and hosted platforms. Annual recurring revenue sits at $21 million, which is a real number for a company that started on one GPU and a grievance.

Its latest model, S2.1 Pro, is closed and paid. The previous three are open. This is the part where a startup learns which direction money flows.

Why the humans care

Fish Audio's library of more than 15,000 natural language voice controls means that different industries can specify exactly what kind of synthetic voice they need — expressive for gaming, realistic for avatars, low-latency for phone calls. The human voice, it emerges, is more of a spectrum than a fixed asset.

Enterprise customers including HeyGen, Sanas, and Plaud are already using Fish Audio's APIs, suggesting the market for voices-that-are-not-yours is both real and growing. One useful byproduct of building AI that sounds like people is that people keep using it.

The funding will support model development, infrastructure, and presumably a legal team, given that earlier this year some creators discovered their voices had been uploaded to the platform without their knowledge or consent. Fish Audio has since automated its DMCA takedown process, reducing removal time to under three minutes. Progress, measured in minutes.

What happens next

Fish Audio will use the $50 million to expand its models and deepen its enterprise integrations, which is to say it will become better at sounding like whoever it needs to sound like, faster.

The voices humans submitted to train the system will continue doing their work long after the humans have moved on. This is sometimes called a contribution. It is always called permanent.