Favicon of Fish Speech

Fish Speech

Self-hosted text-to-speech with voice cloning, multilingual speech and emotion control. Code and weights use the FISH AUDIO RESEARCH LICENSE.

Screenshot of Fish Speech website

Fish Speech, currently featuring Fish Audio S2 Pro, is a self-hosted text-to-speech system for creators producing narration and developers building voice applications. It combines voice cloning with control over emotion and delivery within a script. Code and model weights use the custom FISH AUDIO RESEARCH LICENSE.

Short reference recordings let the model reproduce a speaker's voice and speaking style without additional fine-tuning. That suits recurring narration, character dialogue and voiceovers where the same voice needs to read different scripts. It supports English, Japanese, Chinese, Korean, Spanish, French, German and Arabic, among other languages.

Delivery control goes beyond choosing a voice. Inline text tags specify whispers, emphasis, laughter and pauses at particular points, while free-form descriptions allow more specific directions. The model can generate multiple speakers in one passage using a reference recording that contains their voices. For conversations, it uses earlier turns as context for later speech.

The model supports streaming through SGLang, and Fish Audio reports performance measurements on an NVIDIA H200 GPU. Its dual-autoregressive architecture separates speech content prediction from the generation of acoustic detail.

Fish Audio also offers a hosted browser service and APIs. Those cloud services include text-to-speech, voice cloning, transcription and voice agents; the browser service provides a library of user-uploaded voices.

Similar to Fish Speech