Whisper WebUI turns audio into transcripts and subtitles through a browser interface that runs on your own machine or a self-hosted server. It's for people captioning videos, transcribing recordings or translating spoken content who want local speech processing. The project is open source under Apache 2.0 and supports Docker and Pinokio.
You can transcribe uploaded files, YouTube content or microphone recordings, then export timed subtitles as SRT or WebVTT, or plain text without timestamps. Whisper can also translate speech in other languages into English. For translating existing subtitle text, the app supports Facebook NLLB models locally and the external DeepL API, which sends text to DeepL for processing.
The default backend is faster-whisper, chosen for faster transcription and lower GPU memory use than the original Whisper implementation. You can also use openai/whisper or insanely-fast-whisper, including compatible fine-tuned models downloaded from Hugging Face. The default setup targets NVIDIA GPUs with CUDA; Intel hardware requires a different dependency setup.
Audio processing includes Silero VAD to detect speech and UVR to separate background music. The pyannote integration labels speakers, though access to its models requires a Hugging Face token and acceptance of their terms. A REST API backend supports use beyond the browser interface, and a Google Colab notebook provides an option for running it on hosted hardware.
Claim this page and we'll verify you by hand. Whisper WebUI gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Whisper WebUI?Promote it
Something wrong or outdated on this page?
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
3.3KUpdated 2 months agoMIT
Docker · Web#Multilingual#Speaker diarization#Voice activity detection
21.8KUpdated 1 week agoMIT
macOS · Windows · Linux#Hugging Face integration#Multilingual#Speaker diarization
13.1KUpdated 2 years agoApache-2.0
macOS · Windows#Batch processing#Hugging Face integration#Multilingual
7.7KUpdated 4 weeks agoMIT
macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
Whisper ASR Webservice turns Whisper speech recognition into a self-hosted API for developers adding transcription to their apps or services. It runs in Docker on your own machine or server, with CPU processing or CUDA GPU acceleration. The Python project is open source under the MIT license.
Buzz transcribes and translates speech on your own computer using OpenAI's Whisper. It's for people who need transcripts or subtitles from recordings, plus live captions from a microphone. Local transcription works offline; the optional OpenAI Whisper API sends audio to a cloud service.
insanely-fast-whisper is a command-line tool for people who want to transcribe audio on their own hardware, with a focus on processing long recordings quickly. It runs OpenAI's Whisper locally on NVIDIA GPUs or Apple Silicon Macs, including support for Windows with CUDA. The project is open source under the Apache 2.0 license.
Vibe is an open source desktop app for people who need transcripts or subtitles without uploading their recordings to a transcription service. It runs on macOS, Windows and Linux under the MIT license. Audio transcription works fully offline, with processing on your own computer.