Favicon of NVIDIA Dynamo

NVIDIA Dynamo

A self-hosted inference framework that coordinates NVIDIA GPU clusters with vLLM, SGLang or TensorRT-LLM and exposes an OpenAI-compatible API.

Screenshot of NVIDIA Dynamo website

NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.

Dynamo separates prompt processing (prefill) from token generation (decode), so each phase can use independently scaled GPU pools. Its router considers worker load and cached context when choosing where to send a request, reducing repeated prompt processing. The planner adjusts GPU resources to meet latency targets as workloads change.

Memory management extends beyond GPUs. The KV Block Manager moves cached model context through CPU memory, SSDs and remote storage, freeing GPU memory for active work. NIXL handles cache transfers between GPUs and storage tiers, while ModelExpress streams model weights between GPUs to shorten replica startup times.

For Kubernetes deployments, Grove schedules related components together and accounts for how hardware connects across hosts and racks. Dynamo exposes an OpenAI-compatible API and supports multimodal models and tool calling through its inference backends. It also serves video generation workloads with FastVideo and SGLang Diffusion. The framework uses Rust for performance and Python for extensibility; its AI Perf tool measures serving performance across the supported engines.

Similar to NVIDIA Dynamo