
NVIDIA DCGM monitors and manages NVIDIA data-center GPUs on your own Linux servers. It's for infrastructure teams running GPU clusters, including those hosting AI workloads, who need to track hardware health, investigate slow jobs and control power use. It supports x86_64 and aarch64 (SBSA) systems.
DCGM collects GPU health, performance and utilization data to help teams distinguish hardware faults from application performance problems. Its low-overhead health checks are designed to run alongside active jobs. Active and passive diagnostics help identify failures, performance degradation and power inefficiencies, while system alerts flag problems that need attention.
Management includes power and clock policies, so teams can govern GPU behavior as well as observe it. DCGM works on its own or alongside cluster management and resource scheduling software. DCGM-Exporter exposes GPU metrics and health data in Kubernetes environments for monitoring and alerting; supported integrations include Prometheus, Bright Cluster Manager and IBM Spectrum LSF.
DCGM has an open-core architecture. Its foundational libraries and building blocks are open source under Apache 2.0, while some diagnostics and tests remain proprietary. Teams building their own management software can use its C and Python APIs or Go bindings.
Claim this page with an email at developer.nvidia.com. NVIDIA DCGM gets the verified badge, and you can upgrade the listing to be featured on localhosted. Proud to be listed? Put our badge on your site.
Want more people to find NVIDIA DCGM?Promote it
Something wrong or outdated on this page?
2.3KUpdated 1 day agoMPL-2.0
macOS · Windows · Linux · Docker#Agent Skills#Batch processing#Multi-user access
dstack is a self-hosted orchestration tool for AI teams managing compute across GPU clouds and their own servers. It puts cluster management, training jobs and model inference behind one interface, so teams can use different providers and accelerators without maintaining a separate workflow for each environment. It's open source under the Mozilla Public License 2.0.
223Updated 16 hours agoApache-2.0
macOS · Linux#GGUF#Git integration#Guardrails
6.8KUpdated 4 weeks agoMIT
Windows · Linux
ROCm is AMD's open-source GPU computing platform for developers running AI training, inference and scientific workloads on their own hardware or servers. It supports selected Linux and Windows configurations on AMD Instinct, Radeon and Ryzen AI devices. Check the version-specific GPU, operating-system, driver and firmware compatibility matrix before installing. It's the software foundation for applications that need AMD GPU acceleration, including local LLM workloads.
5.1KUpdated 20 hours agoApache-2.0
#Batch processing#Distributed execution#LoRA
6.9KUpdated 2 days agoApache-2.0
Docker · Web#Code execution#Git integration#Multi-user access
1KUpdated 7 days ago
#Distributed execution#Hugging Face integration#LoRA
Kaito manages self-hosted LLM inference, fine-tuning, and document retrieval services in a Kubernetes cluster. It's for teams that want to run models on infrastructure they control while reducing the work of sizing GPU resources and managing model deployments. The project is open source under Apache 2.0.
LLMKube is a free, open-source Kubernetes operator for teams and homelab owners running local LLM inference across their own hardware. It manages Linux GPU servers and Apple Silicon Macs together, so a mixed fleet can serve models through the same platform. It uses the Apache 2.0 license.
AIBrix is open-source infrastructure for teams serving large language models on their own Kubernetes clusters. It focuses on the work around inference: directing requests, scaling capacity and managing models across servers. Enterprise infrastructure teams can use its components to build a self-hosted model service. It's licensed under Apache 2.0.
ClearML is an MLOps suite for recording experiments, managing datasets and running ML workloads. Its Apache 2.0 Python SDK connects to a ClearML Server, available as a hosted service or open-source software you deploy yourself. ClearML Agent handles job orchestration and reproducibility.