Cloud Platform as a Service

Cloud Platform as a Service

Join us to learn more from a community of collaborative experts and IBM Cloud product users to share advice and best practices with peers and stay up to date regarding product enhancements, regional user group meetings, webinars, how-to blogs, and other helpful materials.

 View Only

Announcing NVIDIA GPU support in IBM Cloud Monitoring for IBM Kubernetes Service and Red Hat OpenShift on IBM Cloud

By Darrell Schrag posted 06/08/26 04:22 PM

  

IBM Cloud Monitoring is introducing out-of-the-box NVIDIA GPU monitoring for IBM Kubernetes Service (IKS) and RedHat OpenShift on IBM Cloud (ROKS) clusters. Simply using the NVIDIA operator installed on your cluster, you'll start seeing all your NVIDIA GPU metrics in your IBM Cloud Monitoring instance, on a curated dashboard, with alerts ready to be enabled in a couple of clicks.

Why GPU monitoring matters

GPUs are probably the most expensive piece of hardware in many Kubernetes clusters today, and they are also the easiest to waste. A misconfigured workload that pins one GPU at 2% utilization, a thermal issue that quietly throttles training jobs, or an NVLink that silently degrades can all cost real money — and SLAs — before anyone notices.

With this integration you get the FinOps angle (am I using what I am paying for?), the reliability angle (is the hardware healthy?) and the performance angle (where is the bottleneck — compute, memory, PCIe?) in a single place, side by side with the rest of your Kubernetes observability.

The NVIDIA GPU Overview dashboard

Once the GPU Operator is running, IBM Cloud Monitoring picks up the metrics automatically and exposes them through a curated dashboard, "NVIDIA GPU Overview". The dashboard can be scoped by cluster, namespace, workload, and GPU device, and it covers the questions an SRE actually asks:
General info: for every GPU in every cluster, product type, GPU family, memory, CUDA runtime, driver version, and MIG configuration.
GPU state: hottest GPU temperature, power draw, GPU utilization, SM and memory clock speeds.
GPU memory: used VRAM and percentage, per device.
Throughput: video encoder/decoder usage, PCIe and NVLink throughput.

NVIDIA GPU Dashboard

Out-of-the-box alerts

The integration also ships three opinionated alerts, pre-loaded in your alert library and ready to enable in a couple of clicks:
GPU Temperature Too High: fires when a GPU stays above 85 °C, the early warning for thermal throttling or hardware damage.
GPU Memory Exhaustion: fires when VRAM usage stays above 95%, before out-of-memory errors crash your training or inference jobs.
GPU Underutilization: fires when a GPU stays below 5% utilization for 15 minutes, so you can spot idle expensive hardware and either reschedule workloads or scale down.

Get started in IBM Kubernetes Service and Red Hat OpenShift on IBM Cloud

Install the NVIDIA operator following our official documentation. Then, after a few minutes, you'll see the dashboard and the alert templates in your IBM Cloud Monitoring instance.

In this article, you learned how IBM Cloud Monitoring helps you monitor your NVIDIA GPUs across your IKS and ROKS clusters with the new Nvidia GPU Overview dashboard and out-of-the-box alerts. Start using IBM Cloud Monitoring today!

Jesús Samitier - Senior Cloud Solution Architect - Sysdig
Gauthami Pulicharla - Product Manager for Cloud Monitoring
Darrell Schrag - Product Manager for Kubernetes and OpenShift services on IBM Cloud

Blog image created by Flat Icons - Flaticon

0 comments
13 views

Permalink