Player not loading? Watch on YouTube
This tutorial walks through running the 3-billion-parameter Orpheus TTS model in Unsloth Desktop for local AI voice generation. The speaker describes Orpheus as an open source model and favors its expressiveness, while acknowledging that newer models may score better on benchmarks.
The setup covers installing Unsloth Desktop, opening its audio menu, downloading Orpheus, and choosing Generate rather than Transcribe. Orpheus appears in safetensors and GGUF formats. The speaker recommends safetensors for a modern GPU and suggests GGUF otherwise. On a GPU with 16 GB of video memory, reported usage reaches about 6 GB after loading, up from 2.6 GB already in use. This is an observed reading, not a minimum hardware requirement. The speaker also reports an installation failure and slower initial generation.
Voice selection happens inside the text box: put a voice name first, followed by a colon, a space, and the dialogue. The examples use Mia and Zac. Emotion tags can appear wherever the effect should occur; the supplied metadata lists laugh, chuckle, sigh, cough, sniffle, groan, yawn, and gasp.
The demonstrated interface lacks a voice selector and a cloning option, although the speaker says the model supports cloning. Generated speech can have long gaps between sentences, which may need regeneration or editing. The tutorial ends with ejecting the model to release memory.