Insanely Fast Whisper: GPU setup and speed comparison

Learn how to configure Insanely Fast Whisper, adjust batch sizes, and add diarization, with a reported 98-second A100 benchmark for 150 minutes of audio.

Player not loading? Watch on YouTube

Insanely Fast Whisper is a command line tool for on-device audio transcription with OpenAI's Whisper Large v3. The video covers installation with pipx, transcription from a file or URL, and GPU configuration for Nvidia CUDA and Apple Silicon through MPS. The supplied metadata identifies its license as Apache 2.0.

The presenter reports A100 benchmarks for 150 minutes of audio: 31 minutes with standard FP32 Whisper, five minutes with FP16, batching and BetterTransformer, and 98 seconds after adding Flash Attention 2. Distil Large v2 reportedly takes 78 seconds, while faster-whisper takes nine minutes and 23 seconds on the same hardware. These are benchmarks reported by the presenter, rather than independently verified timings or expected results for every supported device.

The setup guidance covers enabling Flash Attention 2 and reducing the default batch size of 24 when GPU memory runs short. For Apple Silicon, the presenter recommends MPS and a batch size around four. The video also notes that distilled models trade some accuracy for speed.

Speaker diarization uses pyannote, with flags to constrain the speaker count. Other options include word or chunk timestamps, automatic language detection, and translation of supported speech languages into English. The explanation connects the reported speed gains to half precision, batched audio chunks and memory-efficient attention.