Favicon of faster-whisper

faster-whisper

An open-source Python library for local speech transcription with Whisper models. It runs on CPUs or NVIDIA GPUs and uses CTranslate2.

faster-whisper is a Python library for people building local speech transcription into their own software. It runs OpenAI's Whisper models through CTranslate2, with faster processing and lower memory use than the original Whisper implementation in the project's comparisons. It runs on a CPU. NVIDIA GPUs are supported too, and the code is open source under the MIT license.

It produces timestamped speech segments and can also mark the timing of individual words. Language detection helps when the audio's language isn't known in advance. For recordings with long pauses, Silero VAD can filter out sections without speech; batch processing helps when throughput matters. Eight-bit quantization gives users another way to reduce memory use on both CPU and GPU.

The library works with Whisper checkpoints such as large-v3 and turbo, as well as Distil-Whisper models including distil-large-v3. It can use compatible converted models from a local directory, including fine-tuned Whisper models. Named models can be downloaded from Hugging Face, while locally stored models can be loaded without a hosted transcription service. Projects such as speaches, WhisperX and WhisperLive use it as a transcription backend.

Similar to faster-whisper