
Fish Speech, currently featuring Fish Audio S2 Pro, is a self-hosted text-to-speech system for creators producing narration and developers building voice applications. It combines voice cloning with control over emotion and delivery within a script. Code and model weights use the custom FISH AUDIO RESEARCH LICENSE.
Short reference recordings let the model reproduce a speaker's voice and speaking style without additional fine-tuning. That suits recurring narration, character dialogue and voiceovers where the same voice needs to read different scripts. It supports English, Japanese, Chinese, Korean, Spanish, French, German and Arabic, among other languages.
Delivery control goes beyond choosing a voice. Inline text tags specify whispers, emphasis, laughter and pauses at particular points, while free-form descriptions allow more specific directions. The model can generate multiple speakers in one passage using a reference recording that contains their voices. For conversations, it uses earlier turns as context for later speech.
The model supports streaming through SGLang, and Fish Audio reports performance measurements on an NVIDIA H200 GPU. Its dual-autoregressive architecture separates speech content prediction from the generation of acoustic detail.
Fish Audio also offers a hosted browser service and APIs. Those cloud services include text-to-speech, voice cloning, transcription and voice agents; the browser service provides a library of user-uploaded voices.
Claim this page with an email at fish.audio. Fish Speech gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Fish Speech?Promote it
Something wrong or outdated on this page?
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
8.4KUpdated 4 months agoApache-2.0
#Batch processing#Hugging Face integration#Multilingual
24.2KUpdated 1 day ago
Windows · Linux · Web#Hugging Face integration#Multilingual#Multimodal input
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
26.6KUpdated 2 months agoMIT
Linux#Multilingual#Voice cloning#Voice conversion
Higgs Audio is a family of text-to-speech models from Boson AI for developers building narration and conversational audio. Higgs TTS 2 can adapt pacing and intonation to the text and generate dialogue with distinct speakers across multiple languages.
IndexTTS, currently IndexTTS-2.5, is a local text-to-speech system that can reproduce a speaker's voice using one reference recording. It's for people creating spoken audio and developers building speech generation into their own applications. Voice identity and emotion have separate controls, so an emotional reference can shape the delivery while a different recording supplies the voice.
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
Chatterbox is an MIT-licensed text-to-speech model family for developers and creators who want to generate speech on their own hardware. You can self-host it on a GPU, including in an air-gapped environment, without an account or API key. Resemble AI also offers separate managed hosting.