LocalVocal adds live speech transcription and translation to OBS for streamers and people recording video. It runs Whisper on your own computer, so local captioning works offline and keeps audio processing on your machine. The plugin is open source under GPL-2.0.
Captions can appear in OBS text sources or travel with an RTMP stream to YouTube or Twitch. For recordings, it can save transcripts as TXT or SRT files and synchronize captions with OBS recording timestamps. Partial transcriptions show text while speech is still being processed, and caption filters can remove or replace selected wording.
Translation can stay local through Whisper's built-in translation or neural translation models using CTranslate2. It also supports cloud translation through DeepL and OpenAI, plus translation with a local LLM. Choosing a cloud provider sends translation requests outside your machine and requires an internet connection.
LocalVocal runs on Windows, Linux and macOS, including Intel and Apple Silicon Macs. A GPU isn't required: whisper.cpp supports CPU processing, with optional acceleration on supported NVIDIA and AMD GPUs or through Apple's Metal backend. Vulkan also supports some integrated GPUs.
You can use external GGML Whisper models, including fine-tuned models from Hugging Face, rather than relying only on the bundled Tiny.en model. That gives you a choice of speech models for the languages you caption.
Claim this page and we'll verify you by hand. LocalVocal gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find LocalVocal?Promote it
Something wrong or outdated on this page?
21.8KUpdated 1 week agoMIT
macOS · Windows · Linux#Hugging Face integration#Multilingual#Speaker diarization
Buzz transcribes and translates speech on your own computer using OpenAI's Whisper. It's for people who need transcripts or subtitles from recordings, plus live captions from a microphone. Local transcription works offline; the optional OpenAI Whisper API sends audio to a cloud service.
1.4KUpdated 13 hours agoAGPL-3.0
macOS · Windows · Linux · Browser Extension#Multilingual#ONNX#OpenAI-compatible API
752Updated 1 week agoGPL-3.0
macOS · Windows · Linux · Docker#LM Studio integration#MLX#Multilingual
12.7KUpdated 4 months agoApache-2.0
macOS · Windows · Linux#Multilingual#Streaming inference
PaddleSpeech is a Python toolkit built on PaddlePaddle for developers and researchers building speech applications on their own machines or servers. It covers speech recognition and synthesis, with streaming systems for both. The project uses the Apache 2.0 license and supports Linux, Windows and macOS, with Linux recommended. It supports CPU execution.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
5.4KUpdated 2 days agoMIT
macOS · Windows · Linux#Batch processing#MCP#Multilingual
Sokuji is a free, open source speech translator for people joining meetings across languages. It sends your translated speech through a virtual microphone, so other participants hear ordinary call audio and don't need to install anything. Their replies appear as translated subtitles on your screen.
TranscriptionSuite is a local speech-to-text app for people transcribing lectures, recording conversations or dictating into other apps. It's open source under GPL-3.0, with desktop apps for Windows, macOS and Linux. Transcription works offline after the initial app and model downloads, keeping audio and transcripts on your own hardware.
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
SmartSub is a free, open-source desktop app for people who subtitle recordings or adapt videos into other languages. It combines local transcription, translation, subtitle editing and AI dubbing on Windows, macOS and Linux. Each stage also works independently.