IBM Cloud Monitoring is introducing out-of-the-box NVIDIA GPU monitoring for IBM Kubernetes Service (IKS) and RedHat OpenShift on IBM Cloud (ROKS) clusters. Simply using the NVIDIA operator installed on your cluster, you'll start seeing all your NVIDIA GPU metrics in your IBM Cloud Monitoring instance, on a curated dashboard, with alerts ready to be enabled in a couple of clicks.
Why GPU monitoring matters
GPUs are probably the most expensive piece of hardware in many Kubernetes clusters today, and they are also the easiest to waste. A misconfigured workload that pins one GPU at 2% utilization, a thermal issue that quietly throttles training jobs, or an NVLink that silently degrades can all cost real money — and SLAs — before anyone notices.
With this integration you get the FinOps angle (am I using what I am paying for?), the reliability angle (is the hardware healthy?) and the performance angle (where is the bottleneck — compute, memory, PCIe?) in a single place, side by side with the rest of your Kubernetes observability.
The NVIDIA GPU Overview dashboard
Once the GPU Operator is running, IBM Cloud Monitoring picks up the metrics automatically and exposes them through a curated dashboard, "NVIDIA GPU Overview". The dashboard can be scoped by cluster, namespace, workload, and GPU device, and it covers the questions an SRE actually asks:
General info: for every GPU in every cluster, product type, GPU family, memory, CUDA runtime, driver version, and MIG configuration.
GPU state: hottest GPU temperature, power draw, GPU utilization, SM and memory clock speeds.
GPU memory: used VRAM and percentage, per device.
Throughput: video encoder/decoder usage, PCIe and NVLink throughput.
The integration also ships three opinionated alerts, pre-loaded in your alert library and ready to enable in a couple of clicks:
GPU Temperature Too High: fires when a GPU stays above 85 °C, the early warning for thermal throttling or hardware damage.
GPU Memory Exhaustion: fires when VRAM usage stays above 95%, before out-of-memory errors crash your training or inference jobs.
GPU Underutilization: fires when a GPU stays below 5% utilization for 15 minutes, so you can spot idle expensive hardware and either reschedule workloads or scale down.
Install the NVIDIA operator following our official documentation. Then, after a few minutes, you'll see the dashboard and the alert templates in your IBM Cloud Monitoring instance.
In this article, you learned how IBM Cloud Monitoring helps you monitor your NVIDIA GPUs across your IKS and ROKS clusters with the new Nvidia GPU Overview dashboard and out-of-the-box alerts. Start using IBM Cloud Monitoring today!
Blog image created by Flat Icons - Flaticon