GPT-SoVITS V2 Pro: cloud training and local voice cloning

Learn to train GPT-SoVITS V2 Pro on a cloud RTX 4090, move the models to your computer, and connect speech generation to an OpenClaw skill.

Player not loading? Watch on YouTube

This English overview is based on the video’s full Chinese captions. This tutorial follows a GPT-SoVITS V2 Pro voice cloning workflow, from downloading the package to generating speech on a local computer. The speaker chooses V2 Pro based on the project author's stated preference at the time of recording; this is not a benchmark against V3 or V4.

For training, the demonstration uses a rented RTX 4090 with 24 GB of VRAM and a prepared V2 Pro image. The speaker reports that training stalled on his M2 Mac mini and suggests a cloud GPU for viewers without suitable hardware. He opens JupyterLab, starts the Web UI, and uploads audio extracted from a seven-minute recording.

The preparation steps cover audio slicing, speech recognition, and transcript correction before fine-tuning. Although the speaker explains why accurate labels matter, he skips correcting them in the demonstration. He trains the SoVITS and GPT models, retries GPT training after its output appears missing, and copies both sets of weights to the local installation.

Local inference requires selecting both models and supplying reference audio with its transcript. The interface advises keeping the reference under ten seconds; the speaker says longer clips worked in his own tests. Finally, he explains how his existing OpenClaw skill calls the running API to generate narration for short videos. The API terminal must stay open because closing it stops the service.