
NVIDIA Dynamo is a self-hosted inference framework for teams serving models across multiple GPUs or server nodes. It coordinates SGLang, TensorRT-LLM and vLLM, adding cluster-level scheduling and request routing above those engines. Its focus is large deployments where GPU capacity, response latency and repeated computation affect serving costs.
Dynamo separates prompt processing (prefill) from token generation (decode), so each phase can use independently scaled GPU pools. Its router considers worker load and cached context when choosing where to send a request, reducing repeated prompt processing. The planner adjusts GPU resources to meet latency targets as workloads change.
Memory management extends beyond GPUs. The KV Block Manager moves cached model context through CPU memory, SSDs and remote storage, freeing GPU memory for active work. NIXL handles cache transfers between GPUs and storage tiers, while ModelExpress streams model weights between GPUs to shorten replica startup times.
For Kubernetes deployments, Grove schedules related components together and accounts for how hardware connects across hosts and racks. Dynamo exposes an OpenAI-compatible API and supports multimodal models and tool calling through its inference backends. It also serves video generation workloads with FastVideo and SGLang Diffusion. The framework uses Rust for performance and Python for extensibility; its AI Perf tool measures serving performance across the supported engines.
Claim this page with an email at developer.nvidia.com. NVIDIA Dynamo gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find NVIDIA Dynamo?Promote it
Something wrong or outdated on this page?
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
1.3KUpdated 1 day agoApache-2.0
Web#LoRA#Multimodal input#Ollama integration
4.7KUpdated 23 hours agoApache-2.0
#Batch processing#Distributed execution#OpenAI-compatible API
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
49.3KUpdated 2 hours agoMIT
macOS · Linux · Docker · Web#Code execution#Human approval#llama.cpp backend
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
KubeAI is an open source Kubernetes operator for teams serving AI models on their own infrastructure or cloud clusters. It manages model servers and scales them with demand, including starting from zero running replicas. It uses the Apache 2.0 license and can run on CPUs, GPUs or TPUs, including in a local Kubernetes cluster.
llm-d is an open-source stack for teams serving large language models on their own Kubernetes clusters. It coordinates model servers such as vLLM and SGLang across multiple machines, with routing and resource management for production traffic. It uses the Apache 2.0 license.
LocalAI runs language models, speech, vision and image generation on hardware you control. It's for developers and teams that want a self-hosted AI server for their apps without sending model requests to a cloud service. Its OpenAI-compatible API works with existing clients, and it also accepts Anthropic, Ollama and ElevenLabs API calls.