Voice-Pro brings transcription, voice cloning and multilingual dubbing into a locally run Gradio web app. It's for podcasters, video creators and developers who want to process recordings and generate speech in one interface. The software is free and open source under GPL-3.0.
Its dubbing studio combines YouTube downloads through yt-dlp with audio extraction, Demucs vocal separation and noise removal. Whisper, Faster-Whisper and Whisper-Timestamped handle speech recognition. Subtitle tools include word-level highlighting, subtitle display within video players and translation of files such as SRT, ASS and SSA. It also supports live speech recognition and translation.
For speech generation, F5-TTS, E2-TTS and CosyVoice provide voice cloning from reference audio without training a dedicated voice model. F5-TTS supports fine-tuned language models, while Fun-CosyVoice3 includes Korean. kokoro and Edge-TTS provide other text-to-speech choices, with controls for speed, volume and pitch.
The app runs on your machine, but parts use online services. Translation uses Google's web endpoint through Deep-Translator by default, and Edge-TTS uses an online speech service. Optional Azure Translator and Azure TTS require your own Azure subscription and credentials. Internet access is required.
Windows with an NVIDIA GPU is the verified platform; macOS and Linux operation is unverified. CPU mode is also available. For NVIDIA use, the listed minimum is 4GB of VRAM, with 8GB preferred. Audio exports include WAV, FLAC and MP3.
Claim this page and we'll verify you by hand. Voice-Pro gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Voice-Pro?Promote it
Something wrong or outdated on this page?
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
5.4KUpdated 2 days agoMIT
macOS · Windows · Linux#Batch processing#MCP#Multilingual
5.6KUpdated 2 weeks agoApache-2.0
macOS · Windows · Linux · Web#Multilingual#OpenAI-compatible API#Resumable workflows
8KUpdated 3 days agoMIT
macOS · Web#Batch processing#MLX#Multilingual
1.4KUpdated 1 month agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
SmartSub is a free, open-source desktop app for people who subtitle recordings or adapt videos into other languages. It combines local transcription, translation, subtitle editing and AI dubbing on Windows, macOS and Linux. Each stage also works independently.
YouDub-webui is a self-hosted video translation and dubbing app for creators and small teams who want to process media on their own hardware. It accepts YouTube and Bilibili links or local video files, then produces translated subtitles, cloned-voice dubbing, or both. English-to-Chinese dubbing for YouTube is its most established workflow; it also supports Chinese-to-English dubbing for Bilibili.
MLX Audio is a Python library for developers building speech applications that run locally on Apple Silicon Macs. It uses Apple's MLX framework to accelerate audio models on M-series chips, covering speech generation, transcription and audio cleanup in one library. It's open source under the MIT license.
SoniTranslate is a local AI video dubbing app for creators and translators who need speech in another language to follow the timing of the original video. Its Gradio browser interface brings transcription, translation and speech generation together, with speaker detection for recordings that contain multiple voices. Local installation is tested on Linux, and it can use an NVIDIA GPU or run in CPU mode.