Player not loading? Watch on YouTube
Uploaded on November 24, 2024, this walkthrough follows the speaker's first attempt to install MaskGCT on Windows and test its speech generation. The speaker performs a sparse checkout of the Amphion repository for its MaskGCT component, then installs and tests that component. The setup uses Python 3.10 and a CUDA-enabled PyTorch build. The speaker identifies the PyTorch version as "2.01" and CUDA as "118" in the transcript. No MaskGCT release version is specified; the installation steps describe that dated environment.
Much of the tutorial covers troubleshooting. The speaker works through DLL and voice-loading errors, edits phonemizer to skip problematic voices, and adds UTF-8 encoding to three file-reading locations. A later inference error clears after he corrects FFmpeg's PATH visibility and reopens the code editor from a terminal. These are fixes demonstrated on his system, rather than universal installation requirements.
The local AI tests use his own voice and other reference recordings. He reports roughly six seconds to generate 14 seconds of audio on a 4090. A longer passage reaches about 22 GB of VRAM in his observed run and produces distorted speech. His suggested duration ceiling is an inference based on F5 TTS, not an established MaskGCT limit.
An angry reference sample gives disappointing results in both programs. The speaker considers their quality similar in this brief comparison and prefers F5 TTS for easier installation. He also points to a Windows-specific MaskGCT repository, but explicitly says he did not test it.