Player not loading? Watch on YouTube
This tutorial walks through training a custom RVC voice model in Applio 3.6.4. The speaker describes Applio as an open source voice cloning toolkit and says exported models can work in other RVC-compatible applications. Installation is a prerequisite; this guide starts in the training interface and points to a separate installation tutorial.
The walkthrough names the model and sets a 40,000 Hz sampling rate before preparing the dataset. Viewers can supply an existing dataset path or use the dataset creator. The example uploads three WAV recordings of the same voice, keeps the advanced settings at their defaults, and enables noise filtering and noise reduction. Audio preprocessing comes next, followed by feature extraction. The speaker notes that preprocessing time depends on the length of the recordings.
For training, the speaker recommends a batch size of two and explains how to choose a save interval and epoch count. These are tutorial recommendations rather than measured hardware requirements. After training, the guide generates an index file and selects the trained model and its index for export.
The final demonstrations use an uploaded audio clip for voice conversion, with pitch controls available in advanced settings, and then select the model for text-to-speech. Live voice-call conversion is reserved for a later video.