Favicon of insanely-fast-whisper

insanely-fast-whisper

A local audio transcription CLI that runs Whisper and Distil-Whisper on NVIDIA GPUs or Apple Silicon Macs. Open source under Apache 2.0.

insanely-fast-whisper is a command-line tool for people who want to transcribe audio on their own hardware, with a focus on processing long recordings quickly. It runs OpenAI's Whisper locally on NVIDIA GPUs or Apple Silicon Macs, including support for Windows with CUDA. The project is open source under the Apache 2.0 license.

Its speed comes from GPU batching and optimizations built around Hugging Face Transformers, Optimum and optional Flash Attention 2. It supports Whisper large-v3 and Distil-Whisper large-v2, along with a choice of pretrained speech recognition checkpoints. The published speed benchmarks use an NVIDIA A100, so they shouldn't be treated as expected performance on a personal computer.

You can transcribe audio from a local file or a URL, translate speech, and let Whisper detect the input language automatically. Transcripts can include timestamps for individual words or larger segments. The tool saves results as JSON, and its output converter supports VTT subtitles and plain text.

For recordings with multiple speakers, it uses Pyannote.audio to identify who spoke when. That feature requires a Hugging Face token and acceptance of the gated Pyannote model terms, and supports either a known speaker count or a minimum and maximum range. Audio transcription runs on-device; the CLI targets GPU hardware rather than CPU-only machines. On Apple Silicon, the MPS backend uses more memory than CUDA.

Similar to insanely-fast-whisper