Favicon of Whisper WebUI

Whisper WebUI

A local speech-to-text browser app with Whisper backends, subtitle translation and speaker labeling. Open source under Apache 2.0, with Docker support.

Whisper WebUI turns audio into transcripts and subtitles through a browser interface that runs on your own machine or a self-hosted server. It's for people captioning videos, transcribing recordings or translating spoken content who want local speech processing. The project is open source under Apache 2.0 and supports Docker and Pinokio.

You can transcribe uploaded files, YouTube content or microphone recordings, then export timed subtitles as SRT or WebVTT, or plain text without timestamps. Whisper can also translate speech in other languages into English. For translating existing subtitle text, the app supports Facebook NLLB models locally and the external DeepL API, which sends text to DeepL for processing.

The default backend is faster-whisper, chosen for faster transcription and lower GPU memory use than the original Whisper implementation. You can also use openai/whisper or insanely-fast-whisper, including compatible fine-tuned models downloaded from Hugging Face. The default setup targets NVIDIA GPUs with CUDA; Intel hardware requires a different dependency setup.

Audio processing includes Silero VAD to detect speech and UVR to separate background music. The pyannote integration labels speakers, though access to its models requires a Hugging Face token and acceptance of their terms. A REST API backend supports use beyond the browser interface, and a Google Colab notebook provides an option for running it on hosted hardware.

Similar to Whisper WebUI