
Pricing
Pricing information has not been added yet.
Overview
Fish Speech is an advanced open-source multilingual text-to-speech system developed by Fish Audio. Its S2 Pro model is designed to generate highly natural, expressive speech across more than 80 languages and is trained on a very large multilingual audio corpus, with a focus on realistic prosody, emotion, speaker variation, and conversational output.
The system uses a Dual-Autoregressive architecture that separates slower semantic generation from faster residual acoustic generation. It supports fine-grained inline control through natural-language tags for effects such as whispering, excitement, anger, pauses, laughter, emphasis, pitch changes, and many other vocal behaviors. The model also supports multi-speaker and multi-turn generation and can be served through modern inference stacks.
Fish Speech is intended for developers and researchers building voice assistants, narration systems, conversational characters, dubbing tools, and other speech-generation applications where expressive control matters. The repository includes installation and serving guidance, model information, and benchmark results, while usage of the code and weights is governed by the project’s research license rather than a conventional permissive software license.
Releases
Release information has not been added yet.
Complete your AI stack
Add specialist tools for research, coding, visual content, video, and automation.
Best alternatives
Compare similar tools in the same category.
