WhisperX is an open source speech-to-text tool for people transcribing interviews, meetings, and long recordings on their own computer. It builds on OpenAI's Whisper to produce transcripts with word-level timestamps and optional speaker labels.
Whisper's timestamps mark segments of speech and can drift from the words being spoken. WhisperX uses wav2vec2 alignment to place timestamps on individual words, which helps when making subtitles or finding a passage in a recording. Its pyannote-audio speaker labeling can separate voices in a conversation, though overlapping speech remains difficult and the labels aren't always accurate.
Transcription uses the faster-whisper backend. Batch processing helps with long audio, while voice activity detection identifies stretches that contain speech. An option to carry context between segments can help preserve punctuation and proper nouns. Word alignment needs a model suited to the recording's language; the project provides defaults for English, French, German, Spanish, and Italian, with other models available through Hugging Face.
It runs on a CPU. macOS can use that route, while Windows and Linux can use CUDA GPU acceleration. The documented large-v2 setup uses less than 8 GB of GPU memory. Speaker labeling requires a Hugging Face access token and acceptance of the pyannote model terms. WhisperX is licensed under BSD 2-Clause.
Claim this page and we'll verify you by hand. WhisperX gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find WhisperX?Promote it
Something wrong or outdated on this page?
21.8KUpdated 1 week agoMIT
macOS · Windows · Linux#Hugging Face integration#Multilingual#Speaker diarization
Buzz transcribes and translates speech on your own computer using OpenAI's Whisper. It's for people who need transcripts or subtitles from recordings, plus live captions from a microphone. Local transcription works offline; the optional OpenAI Whisper API sends audio to a cloud service.
19.2KUpdated 2 days agoGPL-3.0
macOS · Windows · Linux · Docker · Web#Batch processing#Human approval#Multilingual
7.7KUpdated 4 weeks agoMIT
macOS · Windows · Linux#Batch processing#Multilingual#Ollama integration
2.1KUpdated 6 hours agoAGPL-3.0
macOS · Windows · Linux#Agent Skills#MCP#Multilingual
14.4KUpdated 24 hours agoMIT
macOS · Windows · Linux#Batch processing#LM Studio integration#Multilingual
18.5KUpdated 3 days agoApache-2.0
macOS · Windows · Linux · Docker · Web#Batch processing#MLX#Multilingual
pyVideoTrans translates spoken audio into another language and produces a video with translated subtitles and AI dubbing. It's for people adapting videos for audiences in other languages who want control over which parts run locally. It recognizes speech directly, so the original video doesn't need subtitles.
Vibe is an open source desktop app for people who need transcripts or subtitles without uploading their recordings to a transcription service. It runs on macOS, Windows and Linux under the MIT license. Audio transcription works fully offline, with processing on your own computer.
OpenChatCut is a local-first AI video editor for creators who want conversational editing with control over the finished cut. AI changes become editable clips, captions, effects and audio tracks in the same project you can adjust manually. It's a free, open-source ChatCut alternative under AGPL-3.0, with a desktop app for macOS, Windows and Linux.
Subtitle Edit is an MIT-licensed subtitle editor for Windows, macOS and Linux. It's for people creating captions, translating dialogue or fixing subtitles that don't match the video. Its core editing, conversion and video playback work offline on your device, with optional AI tools for transcription and translation.
VideoLingo is a self-hosted video translation app for creators and educators who need bilingual subtitles or dubbed versions of their videos. It brings transcription, translation and subtitle timing into one browser interface, with dubbing as an optional output. The project is open source under Apache 2.0; a separate hosted service offers subtitle translation and dubbing.