
llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.
Its main role is to make better use of a cluster as requests compete for compute and cached context. Cache-aware routing sends requests to servers that already hold useful context, while load balancing accounts for demand. Tiered cache storage extends capacity into CPU memory or disk, which can help with repeated and multi-turn requests.
For large models such as DeepSeek-R1 and GPT-OSS, llm-d can separate prompt processing from token generation and distribute mixture-of-experts computation across accelerators. Supported hardware includes NVIDIA and AMD GPUs, Google TPUs, Intel XPUs and CPUs.
Production controls include autoscaling based on inference signals, traffic flow control and fairness for shared deployments. The stack also provides observability and highly available routing. For workloads that don't need an immediate response, it supports asynchronous processing through OpenAI-compatible Batch APIs.
llm-d is aimed at infrastructure teams that need distributed serving rather than a desktop chat app. Its tested deployment recipes cover common serving patterns, and reproducible benchmarks let teams assess performance on specific models and hardware.
Claim this page with an email at llm-d.ai. llm-d gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find llm-d?Promote it
Something wrong or outdated on this page?
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
8.2KUpdated 20 hours ago
#Distributed execution#Multimodal input#OpenAI-compatible API
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
2.6KUpdated 20 hours agoApache-2.0
Web#OpenAI-compatible API#Prompt caching
vLLM Production Stack is an open source inference stack for teams serving LLMs on their own Kubernetes GPU clusters. It brings request routing and monitoring around vLLM, so applications can move from one serving instance to a distributed deployment without changing their code. It requires a GPU-enabled Kubernetes environment.
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.