Favicon of whisper.cpp

whisper.cpp

An open-source speech-to-text engine that runs Whisper models locally on desktop and mobile, with CPU-only inference and GPU acceleration. MIT licensed.

whisper.cpp runs OpenAI's Whisper speech recognition models on your own hardware, with fully offline transcription once you've downloaded a model. It's for developers building speech-to-text into applications and people who want to transcribe audio locally. Audio can stay on-device rather than going to a cloud transcription service. The project is open source under the MIT license.

Its C/C++ implementation uses the ggml machine learning library and provides a C-style API for embedding speech recognition in other software. It also includes a command-line transcription tool and a self-hosted web server. Voice activity detection identifies speech in audio, and the examples include an offline voice assistant.

Platform support covers macOS, Windows, Linux and FreeBSD, plus iOS, Android and Raspberry Pi. WebAssembly lets it run in a browser, and Docker images support server deployments. It works with Whisper models in ggml format, including English-only models and large-v3-turbo. Integer quantization reduces model memory and disk requirements.

You don't need a GPU. CPU-only inference is supported, while NVIDIA CUDA, AMD ROCm and Vulkan provide GPU acceleration. Apple Silicon can use Metal for GPU inference or Core ML and ANEForge for encoder processing on the Apple Neural Engine. OpenVINO supports encoder acceleration on Intel CPUs and GPUs; compatible AMD Ryzen AI processors can use their NPU through VitisAI.

Similar to whisper.cpp