
Sokuji is a free, open source speech translator for people joining meetings across languages. It sends your translated speech through a virtual microphone, so other participants hear ordinary call audio and don't need to install anything. Their replies appear as translated subtitles on your screen.
The desktop app runs on Windows, macOS and Linux, and works with any app that accepts a microphone. It can also subtitle system audio, including videos, streams and recordings. The Chrome and Edge extension handles voice and captions within supported meeting sites, including Zoom, Google Meet, Microsoft Teams and Discord.
Local AI handles speech recognition, translation and speech output entirely on your device. After downloading models, you can use it offline without an account or API key, and your audio stays on your machine. It can use existing CPUs and integrated graphics. Local model choices include Whisper for recognition, Opus-MT and Qwen for translation, and Piper for speech.
For calls, Sokuji offers voice output with incoming subtitles or text-only translation in both directions. Provider-dependent voice cloning can preserve your voice, while adaptive speech speed helps translated speech keep pace. Movable subtitle overlays and conversation exports support presentations and meeting notes.
Cloud options include OpenAI, Google Gemini and Soniox, plus OpenAI-compatible endpoints. In cloud mode, audio goes directly to the chosen provider. The optional paid Kizuna AI account service manages credentials and billing; the app itself uses the AGPL-3.0 license.
Claim this page with an email at sokuji.kizuna.ai. Sokuji gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Sokuji?Promote it
Something wrong or outdated on this page?
5.4KUpdated 2 days agoMIT
macOS · Windows · Linux#Batch processing#MCP#Multilingual
SmartSub is a free, open-source desktop app for people who subtitle recordings or adapt videos into other languages. It combines local transcription, translation, subtitle editing and AI dubbing on Windows, macOS and Linux. Each stage also works independently.
12.7KUpdated 4 months agoApache-2.0
macOS · Windows · Linux#Multilingual#Streaming inference
PaddleSpeech is a Python toolkit built on PaddlePaddle for developers and researchers building speech applications on their own machines or servers. It covers speech recognition and synthesis, with streaming systems for both. The project uses the Apache 2.0 license and supports Linux, Windows and macOS, with Linux recommended. It supports CPU execution.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
14.4KUpdated 9 hours agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
18.6KUpdated 1 day agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
1.4KUpdated 1 month agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
SoniTranslate is a local AI video dubbing app for creators and translators who need speech in another language to follow the timing of the original video. Its Gradio browser interface brings transcription, translation and speech generation together, with speaker detection for recordings that contain multiple voices. Local installation is tested on Linux, and it can use an NVIDIA GPU or run in CPU mode.