Favicon of vLLM Production Stack

vLLM Production Stack

A self-hosted LLM serving stack built on vLLM, with an OpenAI-compatible API, request routing and GPU cluster monitoring. Licensed under Apache 2.0.

vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.

The stack exposes vLLM's OpenAI-compatible API. Teams can host it on infrastructure they control or deploy it on cloud infrastructure, with deployment guidance for AWS, GCP, Azure and Lambda Labs. Its Apache 2.0 license allows teams to adapt the stack for their own deployments.

The router sends requests to backends running different models and supports model aliases. Round-robin routing distributes requests across instances, while session-based routing helps reuse cached context from earlier requests. Kubernetes service discovery and fault tolerance help the router track available serving instances.

LMCache integration adds KV cache offloading, which stores model context outside the GPU cache for reuse. Together with request routing, this gives teams ways to improve serving performance beyond adding more model instances.

Prometheus and Grafana provide a web dashboard for instance health and request load. Teams can inspect response latency, time to first token, queued and active requests, plus GPU cache usage and cache hit rates. The router also exports metrics for each serving instance, including request throughput and uptime.

Similar to vLLM Production Stack