Favicon of SoniTranslate

SoniTranslate

An open-source AI video dubbing app with a browser interface, local Linux execution, CPU or NVIDIA GPU support, and optional cloud APIs.

SoniTranslate is a local AI video dubbing app for creators and translators who need speech in another language to follow the timing of the original video. Its Gradio browser interface brings transcription, translation and speech generation together, with speaker detection for recordings that contain multiple voices. Local installation is tested on Linux, and it can use an NVIDIA GPU or run in CPU mode.

Whisper handles transcription, while speech options include Piper, Coqui XTTS, BARK and Facebook MMS. XTTS can clone a voice from a short recording; OpenVoice and RVC provide other voice imitation options. You can edit translated subtitles and speaker assignments, adjust speech speed and volume, and separate vocals from other audio. Supported languages include English, Japanese, Arabic, Ukrainian and Simplified or Traditional Chinese, though some languages support translation without transcription.

Outputs include dubbed video, separate audio, subtitles by speaker, and video with subtitles alone. It supports SRT and ASS subtitles, including subtitles burned into the video. Batch processing covers multiple files and full YouTube playlists. Document translation also extends to PDF videobooks that display images from the PDF.

The app runs locally, but optional OpenAI transcription, translation and speech generation send content to a cloud API. Colab and an online demo offer hosted alternatives. Pyannote speaker detection requires a Hugging Face account and acceptance of its model terms. The code is open source under Apache 2.0; individual models and weights may carry commercial restrictions.

Similar to SoniTranslate