Player not loading? Watch on YouTube
This Chinese-language tutorial follows the RVC WebUI training and conversion workflow. This tutorial covers RVC 2.3.260718 for AI song covers and live voice conversion. The speaker says older trained models remain compatible and the interface can automatically associate their feature indexes.
Training starts with a recommendation of at least 10 minutes of clean solo audio. For singing, the speaker advises including dry vocals across the intended range: speech alone may produce weaker high notes. The example uses about 12 minutes of material, 40K sampling, V2, 220 epochs and batch size 4. Training takes roughly three and a half hours on the speaker's computer, and the update supports resuming interrupted training.
The cover workflow separates vocals from accompaniment, removes reverb and converts the cleaned vocals with a saved model. Extra separation passes can lose detail. The speaker recommends testing several checkpoints rather than assuming the final one sounds best. In the same-song comparison, they hear little difference between RMVPE and FCPE and attribute high-note quality more to the training material.
Live conversion setup covers microphone selection, pitch, the gender factor and latency settings. On the speaker's RTX 4060, CUDA graph acceleration saves about 0.03 seconds per inference. VB-Cable routes converted audio into OBS; capturing the physical microphone there too would mix both voices. The download guidance specifies a separate NVIDIA 50-series package requiring CUDA 12.8, with another download for AMD and Intel GPUs.