faster-whisper is a Python library for people building local speech transcription into their own software. It runs OpenAI's Whisper models through CTranslate2, with faster processing and lower memory use than the original Whisper implementation in the project's comparisons. It runs on a CPU. NVIDIA GPUs are supported too, and the code is open source under the MIT license.
It produces timestamped speech segments and can also mark the timing of individual words. Language detection helps when the audio's language isn't known in advance. For recordings with long pauses, Silero VAD can filter out sections without speech; batch processing helps when throughput matters. Eight-bit quantization gives users another way to reduce memory use on both CPU and GPU.
The library works with Whisper checkpoints such as large-v3 and turbo, as well as Distil-Whisper models including distil-large-v3. It can use compatible converted models from a local directory, including fine-tuned Whisper models. Named models can be downloaded from Hugging Face, while locally stored models can be loaded without a hosted transcription service. Projects such as speaches, WhisperX and WhisperLive use it as a transcription backend.
Claim this page and we'll verify you by hand. faster-whisper gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find faster-whisper?Promote it
Something wrong or outdated on this page?
54KUpdated 2 days agoMIT
macOS · Windows · Linux · iOS · Android · Docker#Hugging Face integration#Quantization#Streaming inference
whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.
4.7KUpdated 5 days agoMIT
#Batch processing#Multilingual#Quantization
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
2.3KUpdated 1 year agoMIT
Linux#Batch processing#GGUF#Hugging Face integration
8.1KUpdated 2 days agoApache-2.0
#Batch processing#Distributed execution#Hugging Face integration
7.2KUpdated 1 day agoMIT
macOS#Batch processing#Distributed execution#Hugging Face integration
CTranslate2 is an open-source C++ and Python library for developers running Transformer models on their own hardware or servers. It handles translation, text generation, text encoding and speech recognition. Its custom runtime focuses on reducing inference time and memory use compared with general-purpose deep learning frameworks.
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
AutoAWQ is a Python library for developers who want to compress and run LLMs on their own hardware using 4-bit Activation-aware Weight Quantization (AWQ). The project is archived and no longer maintained. It's open source under the MIT license. It installs as a Python package, with optional kernel or Intel CPU dependencies.
LMDeploy is an open-source toolkit for developers serving language and vision-language models on their own hardware. It combines model compression with inference and self-hosted APIs, so teams can use it for batch processing or as the model backend for an application. It uses the Apache 2.0 license.
MLX LM is an open-source Python package for generating text and fine-tuning language models locally on Apple Silicon Macs. Built on MLX, it suits developers and researchers who want to work with models through Python or a terminal, including adapting models to their own tasks. The package uses the MIT license.