GPU Observability for Kubernetes
Get Real-Time Visibility into GPU Usage Across Kubernetes Clusters | NVIDIA Technical Blog

Get Real-Time Visibility into GPU Usage Across Kubernetes Clusters | NVIDIA Technical Blog

5/21/2026

What this post added

This post introduces the open-source GPU Usage Monitor, a project designed to provide real-time visibility into GPU usage across Kubernetes clusters. It details the observability gap in existing Kubernetes monitoring stacks for GPUs and presents a solution that integrates DCGM Exporter, kube-state-metrics, Prometheus, and Grafana via a single Helm chart. The post outlines the architecture, installation process, and key insights provided by the pre-built Grafana dashboards, including GPU allocation trends, compute utilization, memory usage per workload, and running/pending pod counts. It also discusses configuration options for integration with existing Prometheus instances and credential management.

Read the original post ↗