Player not loading? Watch on YouTube
This tutorial covers Minimax H3 voice cloning in Pinokio 8.2.0, alongside the speaker's experience with its DiskSaver plugin. The speaker describes DiskSaver as a way to avoid duplicate model downloads across local AI applications. Their installation froze with an "operation not permitted" error, and a full reset removed installed models and saved configurations. They recommend backing up models before updating.
The speaker estimates the voice model at 20 billion parameters and about 18 GB. Their roughly 50 GB download included both video and voice models plus drivers, rather than the voice model alone. The walkthrough covers installation profiles for limited VRAM, reference audio and text generation without voice cloning.
In the demonstrated interface, reference audio has a 15-second limit and inference defaults to 20 steps. Options include voice LoRAs, skipping inference steps and splitting output by line or paragraph. Saved voice profiles can retain reference audio and output splitting settings.
Results vary in the speaker's tests: a 10-second English clip took 39 minutes, while 15-second generations took three and six minutes. They report a doubled-voice effect, pronunciation errors in English and Spanish, and truncated audio. The final advice covers adjusting VRAM priority when another demanding application, such as DaVinci Resolve, is open.