
vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.
Its main strength is serving capacity. PagedAttention manages the memory used during generation, while continuous batching groups incoming requests to keep hardware busy. Prefix caching and quantization can reduce repeated work and memory use. These capabilities matter when a team needs to serve a model to multiple people or applications.
vLLM supports models from Hugging Face, including Llama, Qwen, Gemma and DeepSeek, as well as multimodal, embedding and classification models. It works with formats and quantization methods such as GGUF, GPTQ and AWQ. Hardware support includes NVIDIA and AMD GPUs, Intel hardware, Apple Silicon and CPUs; Docker is also an option. The OpenAI-compatible API lets applications built for that interface use a model served through vLLM.
Beyond text generation, vLLM supports streaming responses, structured output, tool calling and reasoning parsers. It can serve requests across multiple devices using distributed inference, and its API options include Anthropic Messages and gRPC. Those features make it a fit for teams running their own model service rather than a single-user desktop chat app.
Claim this page with an email at vllm.ai. vLLM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find vLLM?Promote it
Something wrong or outdated on this page?
130KUpdated 37 minutes agoMIT
Web#Code execution#GGUF#Hugging Face integration
llama.cpp runs language models on your own hardware and can serve them from a machine you control. It’s an MIT-licensed, open source inference engine for people building local AI apps, running a private model server, or using a model directly from the command line. It supports vision-language models too.
36.7KUpdated 1 hour agoApache-2.0
#Batch processing#Distributed execution#LoRA
14.7KUpdated 21 hours ago
Docker#Batch processing#Distributed execution#LoRA
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
47.7KUpdated 1 month agoApache-2.0
macOS · Linux · Web#Distributed execution#Hugging Face integration#MLX
3.1KUpdated 3 months agoMIT
macOS · Windows · Linux#Distributed execution#Hugging Face integration#Quantization
SGLang is a self-hosted inference framework for teams that need to serve language and multimodal models on their own hardware. It runs on a single GPU or across distributed clusters and exposes an OpenAI-compatible API. The project is open source under the Apache 2.0 license.
TensorRT-LLM is a library for developers running LLMs on their own NVIDIA GPUs or self-hosted servers. It focuses on inference performance, with support for a single GPU, multiple GPUs, or deployments spread across several machines. Its PyTorch architecture lets teams adapt models and extend the runtime in Python.
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
exo is a local LLM runner that combines your devices into a cluster, letting you use models too large for one machine's memory. It's for people who want to run large models on their own hardware and developers connecting existing AI clients to local inference. It runs on macOS and Linux under the Apache 2.0 license.
Distributed Llama runs a local LLM across several computers, sharing both the computation and the model's memory use. It's for people who want to use their own networked hardware for inference rather than keep the entire workload on one machine. The C++ project is open source under the MIT license.