
llama-cpp-python brings llama.cpp model inference into Python applications and exposes it through a self-hosted OpenAI-compatible server. It's for developers building local AI applications or connecting existing API clients to models on their own hardware. The package is open source under the MIT license.
The Python API handles text generation and chat, with compatibility for LangChain and LlamaIndex. Developers can use its higher-level interface or access llama.cpp's C API directly when they need more control. It loads GGUF model files and can download models from Hugging Face Hub.
The server lets OpenAI-compatible clients send requests to a model running on your machine or server. It supports multiple models and can provide a local Copilot replacement. Function calling lets compatible models request tools, while vision support lets them process images alongside text. Supported image-capable models include LLaVA, moondream2 and qwen2.5-vl; multimodal models also support tool calling and JSON output.
It runs on Linux, Windows and macOS. CPU-only inference is supported, alongside NVIDIA GPU acceleration through CUDA, AMD GPUs through HIP or ROCm, and Apple Silicon acceleration through Metal. Vulkan and SYCL backends are also available. Model inference runs on the hardware hosting the library or server; downloading models from Hugging Face uses an external service.
Claim this page and we'll verify you by hand. llama-cpp-python gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find llama-cpp-python?Promote it
Something wrong or outdated on this page?
qualcomm/GenieXInference Libraries and Bindings
macOS · Windows · Linux#GGUF#Hugging Face integration#llama.cpp backend
Nexa SDK is an on-device AI inference framework for developers building applications that process text, images or audio on users' hardware. It runs models locally across CPUs, GPUs and NPUs, with a shared interface for different backends. Its scope includes language and vision models, speech recognition, speech synthesis and image generation.
1.9KUpdated 3 weeks agoAGPL-3.0
macOS · Windows · Linux · Docker#Batch processing#Distributed execution#Hugging Face integration
77.4KUpdated 1 year agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#OpenAI-compatible API
3.5KUpdated 19 hours agoApache-2.0
macOS · Windows · Linux · iOS · Android · Web#Agent Skills#Hugging Face integration#Multimodal input
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
2.2KUpdated 3 days agoMIT
macOS · Windows · Linux#Batch processing#GGUF#Guardrails
Sonar is a self-hosted inference engine for developers and teams serving Hugging Face-compatible language and multimodal models on their own hardware. Based on vLLM, it adds model and quantization formats, sampling methods, and deployment features. It's open source under AGPL-3.0.
GPT4All is a local AI chatbot for people who want to run language models on their own desktop or laptop and keep conversations on their machine. Its LocalDocs feature lets you ask questions about your own documents without sending them to a cloud service. It suits developers, teams and individuals who want control over their models and data.
LiteRT is Google's open-source framework for developers building AI into apps that run on users' own devices. It succeeds TensorFlow Lite and covers model conversion, optimization and local inference. It's licensed under Apache 2.0.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
node-llama-cpp is an open source library for developers adding local LLM inference to JavaScript and TypeScript applications. It connects Node.js, Bun and Electron to llama.cpp, running GGUF models on your own machine. Its MIT license allows use in commercial projects.