llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
The project uses GGUF models, including models available through Hugging Face. Quantized models reduce memory use, which can make larger models practical on limited hardware. llama.cpp runs on CPUs and can use NVIDIA, AMD, and Apple Silicon GPUs. It can also split work between the CPU and GPU when a model exceeds available video memory.
There’s a built-in web UI for interacting with models and a server with an OpenAI-compatible API for connecting other software. The same engine can run locally or on a server in the cloud; local inference runs on your machine. Built in C and C++ on the ggml library, llama.cpp gives developers a direct way to work with models without requiring a separate hosted inference service.
Claim this page and we'll verify you by hand. llama.cpp gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find llama.cpp?Promote it
Something wrong or outdated on this page?
93KUpdated 1 hour agoApache-2.0
macOS · Docker#Batch processing#Distributed execution#GGUF
vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.
182KUpdated 16 hours agoMIT
macOS · Windows · Linux · Docker#GGUF#llama.cpp backend#Multimodal input
11.9KUpdated 4 days agoAGPL-3.0
macOS · Windows · Linux · Android · Docker · Web#GGUF#Hugging Face integration#llama.cpp backend
7.7KUpdated 5 days agoMIT
macOS · Windows · Linux · Docker · Web#Code execution#GGUF#Hugging Face integration
544Updated 14 hours agoMIT
macOS · iOS · Web#Code execution#Distributed execution#Hugging Face integration
12.5KUpdated 4 months agoApache-2.0
Docker · Web#Hugging Face integration#OpenAI-compatible API
Ollama runs language models on your own computer or server. It provides a command-line runner and a local API for people building AI applications or connecting existing tools to models they host themselves. The software is distributed under the MIT license.
KoboldCpp pairs local model inference with a browser interface built for chat, creative writing and roleplay. A fork of llama.cpp, it bundles KoboldAI Lite with tools for keeping character details and story context alongside your conversations. It's open source under AGPL-3.0.
mistral.rs is an open source inference engine for running models on your own computer or self-hosted server. It's for developers building AI applications and people who want local chat, multimodal models and agent tools in the same runtime. The Rust project uses the MIT license.
Pooled runs a single open model across browser tabs on laptops, desktops and phones, combining their memory when the model won't fit on one device. It's for people who want local AI chat or a coding assistant using hardware they already have. It's open source under the MIT license and requires no account or per-device installation.
OpenLLM is a self-hosted LLM server for developers who want to connect their applications to models running on their own hardware or servers. Its OpenAI-compatible API works with clients built for that interface, including the OpenAI Python client and LlamaIndex. The project is open source under the Apache License 2.0.