VoiceCraft can generate speech or edit an existing recording using just a few seconds of a speaker's voice as reference. It runs locally and suits speech researchers, developers adding voice generation to their projects, and people who want to work with audio on their own hardware.
Its zero-shot approach means a speaker doesn't need a separately trained voice model. The same system handles text-to-speech and speech editing, with examples drawn from audiobooks, internet videos and podcasts. Underneath, a neural codec language model fills in audio tokens to produce the requested speech.
The Gradio interface includes separate generation and editing modes. A long-text mode handles longer passages in parts, while the smart transcript feature lets you supply only the text you want to generate. Standalone scripts support integration into other projects, and the project provides training and fine-tuning code for custom datasets.
Local use supports a Gradio browser interface or Docker, with the Docker route tested on Linux and Windows using NVIDIA GPUs. Google Colab, Hugging Face Spaces and Replicate provide cloud-based ways to try it; those sessions run outside your machine. Pretrained VoiceCraft weights are available through Hugging Face.
The code uses CC BY-NC-SA 4.0, which restricts commercial use, and the weights use the Coqui Public Model License. The project prohibits generating or editing someone's speech without their consent.
Claim this page and we'll verify you by hand. VoiceCraft gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find VoiceCraft?Promote it
Something wrong or outdated on this page?
2.3KUpdated 4 months agoMPL-2.0
macOS · Windows · Linux · Docker#Multilingual#Streaming inference#Voice cloning
XTTS v2 generates speech from text using a reference voice recording or a preset speaker. It runs locally through Coqui TTS and suits developers building speech into apps, as well as researchers who want to fine-tune a speech model on their own hardware.
15.3KUpdated 1 week agoMIT
Docker · Web#Multilingual#Voice cloning
6.4KUpdated 3 years agoMIT
Windows#Hugging Face integration#Multilingual#Voice cloning
StyleTTS 2 is an open-source text-to-speech model for developers and speech researchers who want to generate expressive speech on their own hardware. It can choose a speaking style from the text without a reference recording, while its multispeaker model uses reference audio to reproduce a speaker's voice and delivery. The Python code uses PyTorch and carries the MIT license.
6.3KUpdated 10 months agoApache-2.0
#Hugging Face integration#llama.cpp backend#LoRA
1KUpdated 2 months agoMIT
macOS · Windows · Linux · Docker · Web#Batch processing#Hugging Face integration#LM Studio integration
23.8KUpdated 4 months agoApache-2.0
Linux · Docker · Web#Hugging Face integration#Multilingual#Streaming inference
F5-TTS is a local text-to-speech system that uses a reference recording to generate new speech in that voice without training a separate model for each speaker. It's for developers, speech researchers, and creators who want to generate voices on their own hardware. Its Python code uses MIT, while pretrained models use the noncommercial CC-BY-NC license.
Orpheus TTS is an open-source text-to-speech system for developers building voice applications or adapting speech models to their own recordings. It runs locally and uses a Llama backbone to generate speech with control over emotion and intonation. The code uses the Apache 2.0 license.
Alexandria Audiobook is a local AI audiobook generator for people who want separate narrator and character voices, with control over individual lines. It accepts EPUB, text and Markdown books, uses an LLM to identify speakers and delivery directions, and generates speech with a built-in Qwen3-TTS engine. It's open source under the MIT license.
CosyVoice is a local text-to-speech system for developers and researchers who want to generate speech in a reference speaker's voice, including in another language. Its zero-shot voice cloning doesn't require training a separate model for each speaker. You can run it on your own hardware or deploy it as a self-hosted service.