YouDub-webui is a self-hosted video translation and dubbing app for creators and small teams who want to process media on their own hardware. It accepts YouTube and Bilibili links or local video files, then produces translated subtitles, cloned-voice dubbing, or both. English-to-Chinese dubbing for YouTube is its most established workflow; it also supports Chinese-to-English dubbing for Bilibili.
The app uses Whisper for speech recognition, Demucs to separate speech from background audio, and VoxCPM2 for voice generation. Dubbed videos retain the source music and sound effects. You can also keep the original audio and add burned-in subtitles, or export dubbing without visible subtitles. For local files, an already translated SRT subtitle file lets you skip transcription and automatic translation.
Media processing runs locally. Translation uses an OpenAI-compatible Chat Completions API, which can point to a local service or a cloud provider. A cloud provider receives the text sent for translation. Downloading online videos and fetching models need internet access; task records, cookies, logs and generated videos stay on your machine.
It's open source under Apache 2.0 and runs on Windows, Linux, WSL2 and macOS. An NVIDIA CUDA GPU is recommended for full video processing. CPU processing is possible for parts of the workflow, but transcription, audio separation and speech generation can be very slow.
The browser interface shows task history, progress and errors, with playback and MP4 downloads for finished videos. Failed jobs can resume at the failed stage and reuse completed work.
Claim this page and we'll verify you by hand. YouDub-webui gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find YouDub-webui?Promote it
Something wrong or outdated on this page?
5.4KUpdated 1 day agoMIT
macOS · Windows · Linux#Batch processing#MCP#Multilingual
SmartSub is a free, open-source desktop app for people who subtitle recordings or adapt videos into other languages. It combines local transcription, translation, subtitle editing and AI dubbing on Windows, macOS and Linux. Each stage also works independently.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
13KUpdated 3 months agoGPL-3.0
macOS · Windows · Linux · Web#Hugging Face integration#Multilingual#Quantization
1.4KUpdated 1 month agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#Multilingual
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
1.4KUpdated 11 hours agoAGPL-3.0
macOS · Windows · Linux · Browser Extension#Multilingual#ONNX#OpenAI-compatible API
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
Voice-Pro brings transcription, voice cloning and multilingual dubbing into a locally run Gradio web app. It's for podcasters, video creators and developers who want to process recordings and generate speech in one interface. The software is free and open source under GPL-3.0.
SoniTranslate is a local AI video dubbing app for creators and translators who need speech in another language to follow the timing of the original video. Its Gradio browser interface brings transcription, translation and speech generation together, with speaker detection for recordings that contain multiple voices. Local installation is tested on Linux, and it can use an NVIDIA GPU or run in CPU mode.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.
Sokuji is a free, open source speech translator for people joining meetings across languages. It sends your translated speech through a virtual microphone, so other participants hear ordinary call audio and don't need to install anything. Their replies appear as translated subtitles on your screen.