Monitoring and observability are critical components of any software deployment. IBM Cloud Pak for AIOps provides 14 comprehensive Grafana dashboards that give you deep visibility into your environment’s health, performance, and usage patterns. In this guide, we’ll explore all available dashboards, from high-level overviews to component-specific insights.
Prerequisites
These dashboards belong to an optional monitoring stack that can be enabled post-installation. Both OpenShift and Linux based deployments are supported. Follow the instructions below to enable the monitoring stack and set up a Grafana instance:
All the dashboards below link to their 4.13.0 version, make sure to use the version corresponding to your instance of AIOps.
1. Prometheus Monitoring Stack Dashboard
This dashboard closes the monitoring loop by monitoring the pods that are responsible for collecting metrics from the entire cluster. Making sure these pods and their associated Persistent Volume Claims (PVCs) are healthy ensures a fully functional monitoring stack. This dashboard is critical for ensuring your monitoring infrastructure has adequate CPU, memory, and storage resources.
Key Features:
- Prometheus pod resource utilization (CPU and Memory)
- PVC usage and capacity monitoring
Press enter or click to view image in full size
Prometheus Monitoring Stack Dashboard
High-Level Dashboards
These three dashboards provide high-level visibility into your entire AIOps environment. It is designed for quick health checks and understanding overall system behavior.
2. Top Level Dashboard
The Top Level dashboard is your starting point for monitoring the entire Cloud Pak for AIOps environment. It provides a top-down view to quickly triage issues as well as provides visibility into product-level volume.
The dashboard is comprised of 36 health panels, followed by 6 usage panels at the bottom. The health panels will display green, yellow, or red indicating healthy, warning, or critical states respectively. Yellow or red components should be examined further in the component-specific dashboards.
Key Features:
- Cluster and namespace level resource utilization
- At-a-glance component level health
- Event, alert, incident, topology, metric, and log volume
Press enter or click to view image in full size
Top Level Dashboard
3. Health Dashboard
The Health dashboard focuses on cluster-level resource utilization, node statuses, storage consumption, and namespace-level pod health.
Key Features:
- Cluster level CPU, memory, and filesystem usage
- Node states
- PVC health and utilization
- Pod states
Press enter or click to view image in full size
Health Dashboard
4. Usage Dashboard
The Usage dashboard provides visibility into the traffic flowing through your environment and validates that your Cloud Pak for AIOps is correctly sized to handle the load.
Key Metrics Tracked:
- Events ingested (rate & total)
- Alerts created/updated (rate & total)
- Incidents created/updated (rate & total)
- Logs ingested (rate & total)
- Runbooks executed (rate & total)
- Tickets (ServiceNow, Github, & Jira) created (rate & total)
- ChatOps notifications (rate & total)
Press enter or click to view image in full size
Usage Dashboard
Component-Specific Dashboards
These seven dashboards provide deep visibility into specific components and subsystems of Cloud Pak for AIOps, enabling targeted troubleshooting and performance optimization. Along with unique component metrics, every component dashboard also contains status bars to quickly determine if there is a pod stability or resource issue.
Pod uptime status bars
Pod failures can lead to service disruptions and impact user experience. The “Pod Uptime” status bars will help you quickly identify any pod instability in a component.
Press enter or click to view image in full size
Pod Stability panels present on every component dashboard
The “Lowest Pod Uptimes” status bar will display component pods that have been running for the shortest duration. The color of each panel will correspond to how long the pod has been running:
- 0–1 hour: red
- 1–12 hours: orange
- 12–24 hours: yellow
- 24+ hours: green
If any panels are not green it’s possible that the component is experiencing issues. Debug the pod stability further with the “Component Pods Uptime” graph to see if there have been any uptime gaps or crashback loops.
Resource utilization status bars
Underallocated CPU or memory can lead to slow performance or even crashes. The “Resource Utilization” status bars will help you quickly identify any resource issues in a component.
Press enter or click to view image in full size
Using the status bars is the quickest way to identify pod stability and resource issues affecting a component. Each component dashboard will also contain unique metrics to paint a fuller picture of that component’s performance and health.
5. Analytics Dashboard
The Analytics dashboard monitors the analytical capabilities that derive actionable insights from alerts and events. This is where the “intelligence” in AIOps comes to life.
Key Features:
- AI/ML model health metrics
- Anomaly detection effectiveness
- Pattern recognition analytics
Press enter or click to view image in full size
Analytics Dashboard
6. Event Integrations Dashboard
This dashboard provides visibility into the health and performance of event integrations, showing how external event sources connect to and feed data into your AIOps environment. It focuses on 3 main event integrations: Netcool, Instana, and WebHook.
Key Features:
- Integration health status
- Event ingestion rates per integration
- Connection stability metrics
Press enter or click to view image in full size
Event Integrations Dashboard
7. Event Processing Dashboard
The Event Processing dashboard focuses on 3 major components responsible for processing incoming events: policy automation, datalayer, and runbook automation.
Key Features:
- Events/alerts/incidents totals and rates
- Database performance metrics
- Runbook automation performance
Press enter or click to view image in full size
Event Processing Dashboard
8. Log Integrations Dashboard
This dashboard focuses on the components that handle log ingestion, crucial for log anomaly detection and analysis.
Key Features:
- Log ingestion rates and totals per integration
Press enter or click to view image in full size
Log Integrations Dashboard
9. Metric Integrations Dashboard
The Metric Integrations dashboard provides visibility into components handling metric ingestion, essential for metric anomaly detection.
Key Features:
- Metric ingestion rates and totals per integration
Press enter or click to view image in full size
Metric Integrations Dashboard
10. Topology Dashboard
The Topology dashboard monitors components responsible for ingesting and composing topological resources and their relationships, providing the foundation for topology-based insights.
Key Features:
- Resource ingestion rates and totals per observer
- Request processing and lag metrics
- Database request rates and totals per service
Press enter or click to view image in full size
Topology Dashboard
11. Web UI Dashboard
The Web UI dashboard provides visibility into the components that power the Cloud Pak for AIOps web interface, helping you monitor user experience and interface performance.
Key Features:
- Request rates and totals
- Response status codes
Press enter or click to view image in full size
Web UI Dashboard
Community Dashboards
In addition to the IBM-provided dashboards, three community-maintained dashboards offer deeper visibility into the underlying infrastructure components. These dashboards are not created by nor directly affiliated with IBM, so use them with appropriate caution.
12. Community Flink Dashboard
Apache Flink is a key stream processing framework used within Cloud Pak for AIOps. This community dashboard provides detailed metrics on Flink’s performance and health.
Press enter or click to view image in full size
Community Flink Dashboard
13. Community Kafka Dashboard
Kafka is the messaging backbone of Cloud Pak for AIOps. This Strimzi-based dashboard provides comprehensive Kafka cluster monitoring.
Press enter or click to view image in full size
Community Kafka Dashboard
14. Community Postgres Dashboard
PostgreSQL serves as a crucial database for Cloud Pak for AIOps. This CloudNativePG dashboard provides detailed database metrics.
Press enter or click to view image in full size
Community Postgres Dashboard
Best Practices
To get the most out of these dashboards:
1. Start with the Prometheus Monitoring Stack dashboard to ensure your monitoring infrastructure is healthy
2. Use the high-level dashboards (Top Level, Health, Usage) for daily monitoring and quick health checks
3. Drill down to component dashboards when investigating specific issues
4. Set up alerts on critical metrics to enable proactive monitoring
5. Regularly review usage patterns to optimize resource allocation
6. Keep dashboards updated as you upgrade Cloud Pak for AIOps versions
7. Customize the dashboards to suit your specific needs. Panels can be created, deleted, resized, and reordered to suit your preferences, do whatever is best for your workflow. Backup any modified versions when updating to a new version of a dashboard.
Conclusion
These 14 Grafana dashboards provide comprehensive observability across your entire Cloud Pak for AIOps deployment. From high-level health checks to deep component analysis, you have all the tools needed to maintain a healthy, performant AIOps environment.