Favicon of Whisper ASR Webservice

Whisper ASR Webservice

Self-hosted speech-to-text API that runs in Docker on CPU or CUDA GPUs, with Whisper, Faster Whisper and WhisperX. Open source under MIT.

Whisper ASR Webservice turns Whisper speech recognition into a self-hosted API for developers adding transcription to their apps or services. It runs in Docker on your own machine or server, with CPU processing or CUDA GPU acceleration. The Python project is open source under the MIT license.

You can choose OpenAI Whisper, Faster Whisper or WhisperX as the speech recognition engine. Supported model sizes include tiny, base, small, medium and large-v3, so the service gives you a choice of models alongside the choice of engine. It handles multilingual transcription, speech translation and language identification.

Transcripts can include word-level timestamps, while WhisperX adds speaker diarization to distinguish speakers in a recording. Voice activity detection filters out sections without speech. FFmpeg support lets the service accept a broad range of audio and video formats.

The output fits several uses: plain text for a transcript, JSON or TSV for further processing, and VTT or SRT for subtitles. Its REST API includes Swagger documentation and a browser interface for trying requests, which helps developers assess it as a component in an existing application. Model loading and unloading are configurable, and a persistent model cache avoids repeated downloads when the container starts.

Similar to Whisper ASR Webservice