Favicon of WhisperX

WhisperX

Open source speech-to-text software that runs locally, aligns transcripts word by word, and can label speakers.

WhisperX is an open source speech-to-text tool for people transcribing interviews, meetings, and long recordings on their own computer. It builds on OpenAI's Whisper to produce transcripts with word-level timestamps and optional speaker labels.

Whisper's timestamps mark segments of speech and can drift from the words being spoken. WhisperX uses wav2vec2 alignment to place timestamps on individual words, which helps when making subtitles or finding a passage in a recording. Its pyannote-audio speaker labeling can separate voices in a conversation, though overlapping speech remains difficult and the labels aren't always accurate.

Transcription uses the faster-whisper backend. Batch processing helps with long audio, while voice activity detection identifies stretches that contain speech. An option to carry context between segments can help preserve punctuation and proper nouns. Word alignment needs a model suited to the recording's language; the project provides defaults for English, French, German, Spanish, and Italian, with other models available through Hugging Face.

It runs on a CPU. macOS can use that route, while Windows and Linux can use CUDA GPU acceleration. The documented large-v2 setup uses less than 8 GB of GPU memory. Speaker labeling requires a Hugging Face access token and acceptance of the pyannote model terms. WhisperX is licensed under BSD 2-Clause.

Similar to WhisperX