Silicon Valley Startup Fish Audio Secures $50M Seed Round After Hitting $21M Revenue
Palo Alto-based Fish Audio raised $50 million to expand its AI voice synthesis models for developers and creators. The firm currently reports $21 million in annual recurring revenue and serves over 8 million users through open-source and hosted platforms.
Scaling Synthetic Speech for Global Industry
Fish Audio, a Palo Alto startup founded by former NVIDIA researcher Shijia Liao, recently finalized a $50 million seed funding round led by Coreline Ventures and Capital Today. The investment highlights the company's rapid transition from an open-source side project to a commercial entity generating $21 million in annual recurring revenue. Other participants in the financing included 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0.
The company maintains a library of over 15,000 linguistic controls designed to offer high levels of expressiveness and steerability. Current clients include major AI platforms such as HeyGen, Sanas, and Plaud, who utilize the technology for avatars, gaming characters, and customer service applications. CEO Rissa Cao indicates that while the firm initially operated efficiently with creator-focused plans, the need for advanced model development and enterprise expansion drove the decision to seek external capital.
Addressing Rights and Content Safety
Despite its technical successes, Fish Audio has faced scrutiny regarding unauthorized voice uploads. To mitigate these concerns, the startup implemented an automated system to replace its previous manual DMCA process. The system allows creators to remove infringing content within three minutes by providing a voice sample or contract. Oskue Honda, partner at Coreline Ventures, emphasized that long-term success depends on establishing verified voice ownership and clear licensing terms.
Key operational and technical data for Fish Audio include:
- Model Portfolio: Five releases so far, including four speech generation engines and one speech-to-text model.
- Market Traction: Over 8 million users and 31,000 stars on the Fish Speech GitHub repository.
- Revenue Model: Paid API access for the flagship S2.1 Pro model, alongside tiered monthly subscriptions for voice cloning and generation minutes.
Future Roadmap and Competitive Strategy
Competing in a dense sector against firms like ElevenLabs and Cartesia, Fish Audio plans to launch an audio understanding model and a speech-to-speech engine later this year. According to Rico Mallozzi of 359 Capital, the startup's ability to produce state-of-the-art results with limited resources demonstrates a technical efficiency that narrows the gap between synthetic audio and human performance. The company continues to recruit users to submit voices for training in exchange for compensation, aiming to build a more diverse and nuanced library for enterprise-grade applications.
Source: Tech Crunch
