vLLM and LLM-D: GPU serving and Kubernetes setup

Learn how vLLM serves models and LLM-D routes a Kubernetes fleet, with explanations of KV caching, GPU memory limits and separate prefill/decode pools.

Player not loading? Watch on YouTube

This introduction explains the infrastructure behind self-hosted language model serving, starting with model weights and GPU memory before moving to vLLM and LLM-D. It targets SREs, systems administrators and DevOps engineers; the speaker says no machine learning background is required.

A short PyTorch example uses Llama to explain loading weights and generating text. The course then introduces vLLM as a persistent model server with an OpenAI-compatible API. Although much of the discussion concerns large GPU deployments, the speaker also notes that sufficiently small models can run on a laptop.

The inference explanation separates prefill, which processes the prompt, from decode, which generates tokens. It connects these phases to time to first token and time per output token, then explains KV cache reuse and prefix caching. Batching and sharding address different constraints: serving concurrent requests and fitting a model across GPUs. The speaker emphasizes that model weights and request caches compete for limited GPU memory.

For a fleet of servers, LLM-D introduces routing based on cached work, available memory and queue length. The Kubernetes overview covers separate prefill and decode pools, Helm installation, values.yaml configuration and LeaderWorkerSet groups for sharded models. The speaker presents performance gains as examples; suitable tuning depends on the model, hardware and traffic.