vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.
The stack exposes vLLM's OpenAI-compatible API. Teams can host it on infrastructure they control or deploy it on cloud infrastructure, with deployment guidance for AWS, GCP, Azure and Lambda Labs. Its Apache 2.0 license allows teams to adapt the stack for their own deployments.
The router sends requests to backends running different models and supports model aliases. Round-robin routing distributes requests across instances, while session-based routing helps reuse cached context from earlier requests. Kubernetes service discovery and fault tolerance help the router track available serving instances.
LMCache integration adds KV cache offloading, which stores model context outside the GPU cache for reuse. Together with request routing, this gives teams ways to improve serving performance beyond adding more model instances.
Prometheus and Grafana provide a web dashboard for instance health and request load. Teams can inspect response latency, time to first token, queued and active requests, plus GPU cache usage and cache hit rates. The router also exports metrics for each serving instance, including request throughput and uptime.
Claim this page and we'll verify you by hand. vLLM Production Stack gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find vLLM Production Stack?Promote it
Something wrong or outdated on this page?
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
9.5KUpdated 1 day agoApache-2.0
Docker · Web#Guardrails#MCP#Tool calling
4.7KUpdated 23 hours agoApache-2.0
#Batch processing#Distributed execution#OpenAI-compatible API
2.2KUpdated 20 hours agoApache-2.0
Docker#LLM tracing#MCP#Multi-user access
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.
Higress is a self-hosted AI gateway for developers and teams managing model APIs and the tools their AI agents call. It puts LLM traffic and MCP servers behind a shared entry point, with authentication, traffic controls and monitoring. The open-source edition uses the Apache 2.0 license and runs locally in Docker without registration. Alibaba Cloud also offers a fully managed gateway.
llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.
Agent Router is an open source AI gateway for teams whose agents use both model APIs and MCP tools. It runs on a laptop, a dedicated gateway, or Kubernetes, and gives applications one OpenAI-compatible entry point for cloud providers and self-hosted inference. The project uses the Apache 2.0 license.
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.