Favicon of VoiceCraft

VoiceCraft

Local speech editing and text-to-speech software clones voices from a few seconds of audio, with a Gradio interface and Docker support on Linux and Windows.

VoiceCraft can generate speech or edit an existing recording using just a few seconds of a speaker's voice as reference. It runs locally and suits speech researchers, developers adding voice generation to their projects, and people who want to work with audio on their own hardware.

Its zero-shot approach means a speaker doesn't need a separately trained voice model. The same system handles text-to-speech and speech editing, with examples drawn from audiobooks, internet videos and podcasts. Underneath, a neural codec language model fills in audio tokens to produce the requested speech.

The Gradio interface includes separate generation and editing modes. A long-text mode handles longer passages in parts, while the smart transcript feature lets you supply only the text you want to generate. Standalone scripts support integration into other projects, and the project provides training and fine-tuning code for custom datasets.

Local use supports a Gradio browser interface or Docker, with the Docker route tested on Linux and Windows using NVIDIA GPUs. Google Colab, Hugging Face Spaces and Replicate provide cloud-based ways to try it; those sessions run outside your machine. Pretrained VoiceCraft weights are available through Hugging Face.

The code uses CC BY-NC-SA 4.0, which restricts commercial use, and the weights use the Coqui Public Model License. The project prohibits generating or editing someone's speech without their consent.

Similar to VoiceCraft