Favicon of NVIDIA DCGM

NVIDIA DCGM

GPU management and monitoring software for NVIDIA data-center GPUs on Linux, with Kubernetes telemetry and an Apache 2.0 open-source core.

Screenshot of NVIDIA DCGM website

NVIDIA DCGM monitors and manages NVIDIA data-center GPUs on your own Linux servers. It's for infrastructure teams running GPU clusters, including those hosting AI workloads, who need to track hardware health, investigate slow jobs and control power use. It supports x86_64 and aarch64 (SBSA) systems.

DCGM collects GPU health, performance and utilization data to help teams distinguish hardware faults from application performance problems. Its low-overhead health checks are designed to run alongside active jobs. Active and passive diagnostics help identify failures, performance degradation and power inefficiencies, while system alerts flag problems that need attention.

Management includes power and clock policies, so teams can govern GPU behavior as well as observe it. DCGM works on its own or alongside cluster management and resource scheduling software. DCGM-Exporter exposes GPU metrics and health data in Kubernetes environments for monitoring and alerting; supported integrations include Prometheus, Bright Cluster Manager and IBM Spectrum LSF.

DCGM has an open-core architecture. Its foundational libraries and building blocks are open source under Apache 2.0, while some diagnostics and tests remain proprietary. Teams building their own management software can use its C and Python APIs or Go bindings.

Similar to NVIDIA DCGM