Setting Up Grafana & Prometheus for Multi‐Tenant Environments Storage and Client Clusters - IBM/ibm-spectrum-scale-bridge-for-grafana GitHub Wiki
This guide walks through end-to-end setup of the IBM Storage Scale Bridge for Grafana with Prometheus in a multi-tenant environment - where a storage cluster and one or more client clusters expose metrics independently and Prometheus on the storage node scrapes and merges both.
Overview
┌─────────────────────────────┐
│ pmcollector node │
│ ┌───────────────────────┐ │
│ │ grafana-bridge │ │
│ │ port: 9250 │ │
│ └───────────────────────┘ │
└─────────────┬───────────────┘
│
┌────────────┴────────────┐
▼ ▼
┌─────────────────┐ ┌─────────────────┐
│ Client Cluster │ │ Storage Cluster │
└─────────────────┘ └────────┬────────┘
│
│
▼
┌──────────────┌───────────────┐
│ Prometheus │ Grafana │
│ port: 9090 │ port: 3000 │
└──────────────└───────────────┘
Note: The grafana-bridge must be installed on the pmcollector node of each cluster. Prometheus scrapes metrics from both clusters; Grafana reads from Prometheus as its datasource.
Step 1 - Set Up Grafana Bridge on Both Clusters
Grafana Bridge must be installed, configured, and running on both the storage cluster and the client cluster. Install the bridge only on a pmcollector node of each cluster - the bridge reads performance data collected by the pmcollector daemon.
- For RPM installation -
yum install gpfs.grafana-bridge - For Ansible toolkit installation - Install Grafana Bridge with installation toolkit
- For manual installation - Setup the grafana bridge on a classic IBM Storage Scale cluster
Key configuration files on every node where the bridge is installed:
| File | Purpose |
|---|---|
/etc/grafanabridge/config.ini |
Bridge configuration (ports, authentication, logs, ...) |
/etc/systemd/system/grafana-bridge.service |
systemd unit that manages the bridge process |
Step 2 - Install Grafana on the Storage Cluster
Download grafana.sh from this wiki repository and run it:
curl -O https://raw.githubusercontent.com/IBM/ibm-spectrum-scale-bridge-for-grafana/master/wiki/scripts/grafana.sh
chmod +x grafana.sh
./grafana.sh
The script prompts for a Grafana version (default 12.0.0) and port (default 3000), detects the OS (RHEL/CentOS/Fedora, Ubuntu/Debian, SLES), installs the package, updates grafana.ini, opens the firewall port, and enables the systemd service.
Verify Grafana is running:
curl http://127.0.0.1:3000/api/health
Step 3 - Install Prometheus on the Storage Cluster
Download prometheus.sh from this wiki repository and run it:
curl -O https://raw.githubusercontent.com/IBM/ibm-spectrum-scale-bridge-for-grafana/master/wiki/scripts/prometheus.sh
chmod +x prometheus.sh
./prometheus.sh
Common YAML parsing error
If promtool reports:
FAILED: parsing YAML file /tmp/prometheus.yml: yaml: unmarshal errors:
line 281: cannot unmarshal !!seq into config.plain
The storage.tsdb block uses list syntax instead of map syntax. Fix it:
# Wrong - list syntax
storage:
tsdb:
- out_of_order_time_window: 2d
# Correct - map syntax
storage:
tsdb:
out_of_order_time_window: 2d
Validate after editing:
promtool check config /etc/prometheus/prometheus.yml
# SUCCESS: /etc/prometheus/prometheus.yml is valid prometheus config file syntax
Step 4 - Open Firewall Ports
Open firewall port 3000 (Grafana) & 9090 (Prometheus) on storage cluster unidirectional. Firewall port 9250 on both storage & all client clusters.
Step 5 - Fetch and Merge Prometheus Configurations
Each bridge instance exposes a ready-made Prometheus scrape configuration at /prometheus.yml. Fetch both, then merge them into a single prometheus.yml using the script below.
5.1 Fetch each cluster's scrape config
The bridge exposes a ready-made Prometheus scrape config at /prometheus.yml. You can append a jobname_suffix query parameter to tag jobs with a cluster name or hostname, which helps distinguish targets in Prometheus:
# Storage cluster bridge
curl -k "https://<storage_cluster_ip>:9250/prometheus.yml?jobname_suffix=<clustername_or_hostname>" \
-o /tmp/prom_storage.yml
# Client cluster bridge (also runs on port 9250)
curl -k "https://<client_cluster_ip>:9250/prometheus.yml?jobname_suffix=<clustername_or_hostname>" \
-o /tmp/prom_client.yml
5.2 Run the merge script
Download merge_prometheus_configs.py from this wiki repository:
curl -O https://raw.githubusercontent.com/IBM/ibm-spectrum-scale-bridge-for-grafana/master/wiki/scripts/merge_prometheus_configs.py
chmod +x merge_prometheus_configs.py
Run it (all flags are optional - defaults match the standard layout):
python3 merge_prometheus_configs.py \
--storage-yml /tmp/prom_storage.yml \
--client-yml /tmp/prom_client.yml \
--output /etc/prometheus/prometheus.yml \
--storage-cluster-name storage-cluster \
--client-cluster-name client-cluster
Validate and reload:
promtool check config /etc/prometheus/prometheus.yml && systemctl restart prometheus
Step 6 - Verify Endpoints and Metrics Collection
Bridge endpoints
# Storage bridge
curl -k https://<storage_cluster_ip>:9250/endpoints
curl -k https://<storage_cluster_ip>:9250/metadata/time
# Client bridge
curl -k https://<client_cluster_ip>:9250/endpoints
curl -s -k https://<client_cluster_ip>:9250/metadata/sensormetrics | python3 -m json.tool
Note: Open the Grafana UI at
http://<storage_cluster_ip>:3000, add your Prometheus datasource pointing tohttp://<storage_cluster_ip>:9090and start building dashboards to see all the storage and client metrics in a same dashboard. Use theclusterNamelabel in PromQL to filter by cluster:gpfs_health_status{clusterName="storage-cluster"} gpfs_health_status{clusterName="client-cluster"}