Favicon of llama.cpp

llama.cpp

An open source local LLM engine for GGUF models, with CPU and GPU support, a built-in web UI, and an OpenAI-compatible server.

llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.

The project uses GGUF models, including models available through Hugging Face. Quantized models reduce memory use, which can make larger models practical on limited hardware. llama.cpp runs on CPUs and can use NVIDIA, AMD, and Apple Silicon GPUs. It can also split work between the CPU and GPU when a model exceeds available video memory.

There’s a built-in web UI for interacting with models and a server with an OpenAI-compatible API for connecting other software. The same engine can run locally or on a server in the cloud; local inference runs on your machine. Built in C and C++ on the ggml library, llama.cpp gives developers a direct way to work with models without requiring a separate hosted inference service.

Similar to llama.cpp