RVC setup: install and train a local voice model

Learn RVC's portable installation and voice training workflow, using 10 to 15 second WAV clips, feature extraction, and index training.

Player not loading? Watch on YouTube

This tutorial walks through installing RVC, or retrieval-based voice conversion, and training a custom voice model on your own computer. The speaker describes voice conversion as changing a voice while retaining the original speech and timing. They also discuss vocal separation and live voice changing, but defer the real-time calling setup to a later video.

The installation uses a portable release downloaded from GitHub. The speaker lists CUDA Toolkit 11.8 to 12.8 and points out separate downloads for graphics cards below the Nvidia 50 series and for 50 series cards. They also state that AMD and Intel graphics can be used, though the demonstrated download is for Nvidia. After extracting the archive with 7-Zip, the user launches the WebUI from the extracted folder.

Training starts with a model name and a folder of voice recordings. The speaker recommends WAV clips lasting 10 to 15 seconds, selects two CPU cores, and leaves the other initial settings at their defaults. The workflow then processes the audio and extracts features. For training, the demonstration sets "ep show" to 50 and training to 100, with progress visible in the command window.

After model training finishes, the tutorial trains an index file and refreshes the voice list in the inference tab. This is a local AI setup and training walkthrough; it ends before a worked speech-conversion example or live calling demonstration.