Favicon of llama-cpp-python

llama-cpp-python

Python library for running GGUF models locally through llama.cpp, with a self-hosted OpenAI-compatible server and CPU or GPU support.

Screenshot of llama-cpp-python website

llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.

The Python API handles text generation and chat, with compatibility for LangChain and LlamaIndex. Developers can use its higher-level interface or access llama.cpp's C API directly when they need more control. It loads GGUF model files and can download models from Hugging Face Hub.

The server lets OpenAI-compatible clients send requests to a model running on your machine or server. It supports multiple models and can provide a local Copilot replacement. Function calling lets compatible models request tools, while vision support lets them process images alongside text. Supported image-capable models include LLaVA, moondream2 and qwen2.5-vl; multimodal models also support tool calling and JSON output.

It runs on Linux, Windows and macOS. CPU-only inference is supported, alongside NVIDIA GPU acceleration through CUDA, AMD GPUs through HIP or ROCm, and Apple Silicon acceleration through Metal. Vulkan and SYCL backends are also available. Model inference runs on the hardware hosting the library or server; downloading models from Hugging Face uses an external service.

Similar to llama-cpp-python