Favicon of Orpheus TTS

Orpheus TTS

Open-source text-to-speech built on Llama, with local inference, voice cloning and streaming audio. Uses Apache 2.0; Baseten offers cloud hosting.

Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.

The English model includes named voices such as tara, leah and leo. Tags can cue laughter, sighs, gasps and other vocal expressions within spoken text, so applications can specify delivery alongside the words themselves. Streaming audio supports interactive speech applications without waiting for the whole recording to finish.

Voice cloning uses reference text and audio without requiring a separate fine-tuning stage. The pretrained model accepts those examples as context, though it wasn't explicitly trained for zero-shot cloning; additional reference pairs can improve how reliably it matches a voice. A separate fine-tuned English model targets everyday speech generation.

Local inference uses vLLM through the Python package, and a llama.cpp implementation provides a CPU-only route. Baseten offers cloud inference with real-time streaming, while Colab notebooks provide another hosted way to try the models.

For custom voices or speech styles, the project includes training scripts, data preparation tools and sample datasets. Fine-tuning supports Hugging Face tooling and LoRA. An optional Silent Cipher integration adds watermarks to generated audio.

Similar to Orpheus TTS