Player not loading? Watch on YouTube
Voice-Pro is presented as an open source Python app for voice work, with a Gradio interface running on the user's machine. The overview covers voice cloning, transcription, translation, speech generation and vocal isolation, then explains how the dubbing studio connects them.
The speaker describes F5-TTS, E2-TTS and CosyVoice 2.0 as cloning options that accept a 5-to-15-second voice sample. CosyVoice's first download reportedly includes 9 GB of weights. Transcription choices include Whisper, Faster-Whisper, Whisper timestamped for word timing, and WhisperX for speaker labels. Edge-TTS and Kokoro handle speech generation; the speaker says Kokoro can run on CPU.
The dubbing example starts with a YouTube URL. yt-dlp retrieves the media, Demucs separates vocals, Whisper produces a timestamped script, and Deep Translator translates it. A cloning engine generates the replacement speech, which Voice-Pro places back on the original video timeline.
Setup focuses on Windows: clone the repository, run configure.bat, then start.bat. An Nvidia GPU is strongly recommended. The speaker calls Mac and Linux unverified, citing community reports that they work. Comparisons with ElevenLabs describe the speaker's assessment rather than a controlled quality test. The closing caveat is maintenance: the speaker reports that the original team's active development has paused and describes the code as GPL.