Favicon of Distil-Whisper

Distil-Whisper

Local speech recognition models for English transcription, built on Whisper. Run on CPU or CUDA GPUs with Hugging Face Transformers under the MIT license.

Distil-Whisper is a family of local speech recognition models for developers building English transcription into their apps or services. It reduces Whisper's size and processing time while retaining much of its transcription accuracy. It supports English only.

The models run on your own hardware through Hugging Face Transformers, with CPU and CUDA GPU support. The project is open source under the MIT license, which permits commercial use. Published benchmarks report faster inference and smaller models than Whisper, though accuracy varies by model and audio length. Tests also show resistance to background noise and fewer repeated words and inserted errors than Whisper.

For general transcription, distil-large-v3 works across Whisper libraries and handles both short clips and long recordings. Long-form processing offers a choice between sequential transcription for accuracy and chunked transcription for lower latency on individual files. Segment timestamps are supported. The smaller distil-small.en model targets devices with limited memory, including mobile applications, at a cost in accuracy.

Distil-Whisper can also act as an assistant to the original Whisper model through speculative decoding. This approach preserves Whisper's exact output with faster inference in reported benchmarks. The repository includes training code that developers can adapt to distill Whisper for other languages; the supplied recognition models remain English-only.

Similar to Distil-Whisper