Favicon of vLLM

vLLM

An open source LLM serving engine that runs on your hardware, supports NVIDIA and AMD GPUs, and provides an OpenAI-compatible API.

Screenshot of vLLM website

vLLM is an open source engine for serving large language models on hardware you control. It suits developers and teams that need to handle many requests through an API while making efficient use of memory and compute. It's licensed under Apache 2.0 and can run with GPUs or on a CPU.

Its main strength is serving capacity. PagedAttention manages the memory used during generation, while continuous batching groups incoming requests to keep hardware busy. Prefix caching and quantization can reduce repeated work and memory use. These capabilities matter when a team needs to serve a model to multiple people or applications.

vLLM supports models from Hugging Face, including Llama, Qwen, Gemma and DeepSeek, as well as multimodal, embedding and classification models. It works with formats and quantization methods such as GGUF, GPTQ and AWQ. Hardware support includes NVIDIA and AMD GPUs, Intel hardware, Apple Silicon and CPUs; Docker is also an option. The OpenAI-compatible API lets applications built for that interface use a model served through vLLM.

Beyond text generation, vLLM supports streaming responses, structured output, tool calling and reasoning parsers. It can serve requests across multiple devices using distributed inference, and its API options include Anthropic Messages and gRPC. Those features make it a fit for teams running their own model service rather than a single-user desktop chat app.

Similar to vLLM