Favicon of pyVideoTrans

pyVideoTrans

An open-source video translation and AI dubbing tool for Windows, macOS and Linux, with local offline models or cloud APIs. Licensed under GPL-3.0.

Screenshot of pyVideoTrans website

pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.

The workflow covers speech recognition, translation, voice generation and the finished video. You can pause to proofread at each stage. For conversations, speaker diarization separates speakers, and multi-role dubbing assigns them different voices. Voice cloning works with models such as F5-TTS, CosyVoice and GPT-SoVITS.

Local processing can use Faster-Whisper for transcription, Ollama or M2M100 for offline translation, and locally deployed speech models for dubbing. With local models, processing can stay on your own hardware and work offline. Cloud choices include ChatGPT, DeepSeek, Claude, Gemini, OpenAI and Azure; using those services sends the relevant processing to their APIs.

The GPL-3.0 open-source software runs on Windows, macOS and Linux without an application account. It has a desktop interface, a browser interface for remote or internal network access, and command-line support for batch work and self-hosted servers, including Docker deployment. An NVIDIA GPU is optional and can accelerate processing through CUDA.

Separate tools handle batch audio transcription, SRT subtitle translation, text-to-speech, vocal separation and audio-video alignment.

Similar to pyVideoTrans