Favicon of llm-d

llm-d

An open-source LLM inference stack for self-hosted Kubernetes clusters, with vLLM and SGLang backends and support for GPUs, TPUs, XPUs and CPUs.

Screenshot of llm-d website

llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.

Its main role is to make better use of a cluster as requests compete for compute and cached context. Cache-aware routing sends requests to servers that already hold useful context, while load balancing accounts for demand. Tiered cache storage extends capacity into CPU memory or disk, which can help with repeated and multi-turn requests.

For large models such as DeepSeek-R1 and GPT-OSS, llm-d can separate prompt processing from token generation and distribute mixture-of-experts computation across accelerators. Supported hardware includes NVIDIA and AMD GPUs, Google TPUs, Intel XPUs and CPUs.

Production controls include autoscaling based on inference signals, traffic flow control and fairness for shared deployments. The stack also provides observability and highly available routing. For workloads that don't need an immediate response, it supports asynchronous processing through OpenAI-compatible Batch APIs.

llm-d is aimed at infrastructure teams that need distributed serving rather than a desktop chat app. Its tested deployment recipes cover common serving patterns, and reproducible benchmarks let teams assess performance on specific models and hardware.

Similar to llm-d