Whisper ASR Webservice turns Whisper speech recognition into a self-hosted API for developers adding transcription to their apps or services. It runs in Docker on your own machine or server, with CPU processing or CUDA GPU acceleration. The Python project is open source under the MIT license.
You can choose OpenAI Whisper, Faster Whisper or WhisperX as the speech recognition engine. Supported model sizes include tiny, base, small, medium and large-v3, so the service gives you a choice of models alongside the choice of engine. It handles multilingual transcription, speech translation and language identification.
Transcripts can include word-level timestamps, while WhisperX adds speaker diarization to distinguish speakers in a recording. Voice activity detection filters out sections without speech. FFmpeg support lets the service accept a broad range of audio and video formats.
The output fits several uses: plain text for a transcript, JSON or TSV for further processing, and VTT or SRT for subtitles. Its REST API includes Swagger documentation and a browser interface for trying requests, which helps developers assess it as a component in an existing application. Model loading and unloading are configurable, and a persistent model cache avoids repeated downloads when the container starts.
Claim this page and we'll verify you by hand. Whisper ASR Webservice gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find Whisper ASR Webservice?Promote it
Something wrong or outdated on this page?
2.9KUpdated 9 months agoApache-2.0
Windows · Docker · Web#Hugging Face integration#Multilingual#Speaker diarization
Whisper WebUI turns audio into transcripts and subtitles through a browser interface that runs on your own machine or a self-hosted server. It's for people captioning videos, transcribing recordings or translating spoken content who want local speech processing. The project is open source under Apache 2.0 and supports Docker and Pinokio.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
1.5KUpdated 2 months agoMIT
Docker#Batch processing#Multilingual#OpenAI-compatible API
3.1KUpdated 1 year agoAGPL-3.0
Web#Multilingual#Works offline
Whishper is a self-hosted speech-to-text app for people who need transcripts or translated subtitles from audio and video. Its browser interface brings transcription, translation and subtitle editing together, with all three running on your own machine. It can work offline, so local media doesn't need to go to a cloud transcription service.
3.7KUpdated 5 months agoMIT
Docker#OpenAI-compatible API#Streaming inference
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.
subgen generates subtitles on your own hardware for personal media libraries, including films and shows that don't have usable subtitles available. It's an open source, MIT-licensed Python service that runs in Docker or as a standalone application. Speech recognition runs locally using Whisper models through faster-whisper and stable-ts, with support for CPU processing and NVIDIA GPUs through CUDA.
Speaches is a self-hosted speech server for developers who want transcription, translation and speech generation on their own hardware. Its OpenAI-compatible API lets applications use local speech models through tools and SDKs built for OpenAI's API. The project is open source under the MIT license.