Favicon of Voice-Pro

Voice-Pro

Local AI dubbing and voice cloning app with Whisper transcription and a browser interface. Free under GPL-3.0; requires internet access.

Voice-Pro brings transcription, voice cloning and multilingual dubbing into a locally run Gradio web app. It's for podcasters, video creators and developers who want to process recordings and generate speech in one interface. The software is free and open source under GPL-3.0.

Its dubbing studio combines YouTube downloads through yt-dlp with audio extraction, Demucs vocal separation and noise removal. Whisper, Faster-Whisper and Whisper-Timestamped handle speech recognition. Subtitle tools include word-level highlighting, subtitle display within video players and translation of files such as SRT, ASS and SSA. It also supports live speech recognition and translation.

For speech generation, F5-TTS, E2-TTS and CosyVoice provide voice cloning from reference audio without training a dedicated voice model. F5-TTS supports fine-tuned language models, while Fun-CosyVoice3 includes Korean. kokoro and Edge-TTS provide other text-to-speech choices, with controls for speed, volume and pitch.

The app runs on your machine, but parts use online services. Translation uses Google's web endpoint through Deep-Translator by default, and Edge-TTS uses an online speech service. Optional Azure Translator and Azure TTS require your own Azure subscription and credentials. Internet access is required.

Windows with an NVIDIA GPU is the verified platform; macOS and Linux operation is unverified. CPU mode is also available. For NVIDIA use, the listed minimum is 4GB of VRAM, with 8GB preferred. Audio exports include WAV, FLAC and MP3.

Similar to Voice-Pro