Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
The English model includes named voices such as tara, leah and leo. Tags can cue laughter, sighs, gasps and other vocal expressions within spoken text, so applications can specify delivery alongside the words themselves. Streaming audio supports interactive speech applications without waiting for the whole recording to finish.
Voice cloning uses reference text and audio without requiring a separate fine-tuning stage. The pretrained model accepts those examples as context, though it wasn't explicitly trained for zero-shot cloning; additional reference pairs can improve how reliably it matches a voice. A separate fine-tuned English model targets everyday speech generation.
Local inference uses vLLM through the Python package, and a llama.cpp implementation provides a CPU-only route. Baseten offers cloud inference with real-time streaming, while Colab notebooks provide another hosted way to try the models.
For custom voices or speech styles, the project includes training scripts, data preparation tools and sample datasets. Fine-tuning supports Hugging Face tooling and LoRA. An optional Silent Cipher integration adds watermarks to generated audio.
Claim this page and we'll verify you by hand. Orpheus TTS gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Orpheus TTS?Promote it
Something wrong or outdated on this page?
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
32.9KUpdated 2 weeks ago
#Batch processing#Multilingual#Multimodal input
Fish Speech, currently featuring Fish Audio S2 Pro, is a self-hosted text-to-speech system for creators producing narration and developers building voice applications. It combines voice cloning with control over emotion and delivery within a script. Code and model weights use the custom FISH AUDIO RESEARCH LICENSE.
62.2KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.