Whispering Tiger turns audio on your computer into live transcripts and translations for VRChat and streaming overlays. It's free, open source software under the MIT license, aimed at people who want captions or translated conversations while gaming or broadcasting. Processing stays on your machine, and it works offline once you've downloaded the models.
Its connection to other apps is a central feature. It sends text to VRChat through OSC or to browser overlays through WebSockets, with support for streaming applications such as OBS, vMix and XSplit. It can listen to microphone input and computer audio, so its use extends beyond transcribing your own speech.
Speech recognition options include OpenAI's Whisper, Meta's Seamless M4T and NVIDIA's NeMo Canary. For text translation, it supports NLLB-200 and M2M-100, alongside automatic language identification. CUDA GPU acceleration is available.
It also reads text in game images. EasyOCR handles text recognition, while GOT-OCR 2.0 and Microsoft's Phi-4 Multimodal LLM support handwriting. That gives players a way to translate visible text as well as spoken dialogue.
Text-to-speech can read transcripts and translations aloud. Supported engines include Kokoro TTS and Silero; F5/E2-TTS and Zonos add voice cloning and streamed playback. An RVC plugin can change voices from live audio, audio files or generated speech. A separate native interface manages settings and application updates.
Claim this page and we'll verify you by hand. Whispering Tiger gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Whispering Tiger?Promote it
Something wrong or outdated on this page?
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
5.4KUpdated 2 days agoMIT
macOS · Windows · Linux#Batch processing#MCP#Multilingual
20.3KUpdated 5 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#Multilingual#Voice cloning
14.4KUpdated 9 hours agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
1.4KUpdated 1 month agoApache-2.0
Linux · Web#Batch processing#Hugging Face integration#Multilingual
62.3KUpdated 1 month agoMIT
macOS · Windows · Linux · Docker · Web#Hugging Face integration#Multilingual#Voice activity detection
SmartSub is a free, open-source desktop app for people who subtitle recordings or adapt videos into other languages. It combines local transcription, translation, subtitle editing and AI dubbing on Windows, macOS and Linux. Each stage also works independently.
ebook2audiobook turns non-DRM ebooks into narrated audio with chapters and metadata, for readers who want audio editions of their own books. It runs locally on Windows, macOS and Linux, with Docker support and a browser interface built with Gradio. It's open source under Apache 2.0.
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
SoniTranslate is a local AI video dubbing app for creators and translators who need speech in another language to follow the timing of the original video. Its Gradio browser interface brings transcription, translation and speech generation together, with speaker detection for recordings that contain multiple voices. Local installation is tested on Linux, and it can use an NVIDIA GPU or run in CPU mode.
GPT-SoVITS is a local text-to-speech and voice cloning tool. It can generate speech from a short reference recording or fine-tune a model for a custom voice. The source code uses the MIT license.