Setting Up Grafana & Prometheus for Multi‐Tenant Environments Storage and Client Clusters - IBM/ibm-spectrum-scale-bridge-for-grafana GitHub Wiki

This guide walks through end-to-end setup of the IBM Storage Scale Bridge for Grafana with Prometheus in a multi-tenant environment - where a storage cluster and one or more client clusters expose metrics independently and Prometheus on the storage node scrapes and merges both.


Overview

           ┌─────────────────────────────┐
           │       pmcollector node      │
           │  ┌───────────────────────┐  │
           │  │   grafana-bridge      │  │
           │  │   port: 9250          │  │
           │  └───────────────────────┘  │
           └─────────────┬───────────────┘
                         │ 
            ┌────────────┴────────────┐
            ▼                         ▼
   ┌─────────────────┐       ┌─────────────────┐
   │  Client Cluster │       │ Storage Cluster │
   └─────────────────┘       └────────┬────────┘
                                      │
                                      │
                                      ▼
                       ┌──────────────┌───────────────┐
                       │  Prometheus  │   Grafana     │
                       │  port: 9090  │   port: 3000  │
                       └──────────────└───────────────┘

Note: The grafana-bridge must be installed on the pmcollector node of each cluster. Prometheus scrapes metrics from both clusters; Grafana reads from Prometheus as its datasource.


Step 1 - Set Up Grafana Bridge on Both Clusters

Grafana Bridge must be installed, configured, and running on both the storage cluster and the client cluster. Install the bridge only on a pmcollector node of each cluster - the bridge reads performance data collected by the pmcollector daemon.

Key configuration files on every node where the bridge is installed:

File Purpose
/etc/grafanabridge/config.ini Bridge configuration (ports, authentication, logs, ...)
/etc/systemd/system/grafana-bridge.service systemd unit that manages the bridge process

Step 2 - Install Grafana on the Storage Cluster

Download grafana.sh from this wiki repository and run it:

curl -O https://raw.githubusercontent.com/IBM/ibm-spectrum-scale-bridge-for-grafana/master/wiki/scripts/grafana.sh
chmod +x grafana.sh
./grafana.sh

The script prompts for a Grafana version (default 12.0.0) and port (default 3000), detects the OS (RHEL/CentOS/Fedora, Ubuntu/Debian, SLES), installs the package, updates grafana.ini, opens the firewall port, and enables the systemd service.

Verify Grafana is running:

curl http://127.0.0.1:3000/api/health

Step 3 - Install Prometheus on the Storage Cluster

Download prometheus.sh from this wiki repository and run it:

curl -O https://raw.githubusercontent.com/IBM/ibm-spectrum-scale-bridge-for-grafana/master/wiki/scripts/prometheus.sh
chmod +x prometheus.sh
./prometheus.sh

Common YAML parsing error

If promtool reports:

FAILED: parsing YAML file /tmp/prometheus.yml: yaml: unmarshal errors:
  line 281: cannot unmarshal !!seq into config.plain

The storage.tsdb block uses list syntax instead of map syntax. Fix it:

# Wrong - list syntax
storage:
  tsdb:
    - out_of_order_time_window: 2d

# Correct - map syntax
storage:
  tsdb:
    out_of_order_time_window: 2d

Validate after editing:

promtool check config /etc/prometheus/prometheus.yml
# SUCCESS: /etc/prometheus/prometheus.yml is valid prometheus config file syntax

Step 4 - Open Firewall Ports

Open firewall port 3000 (Grafana) & 9090 (Prometheus) on storage cluster unidirectional. Firewall port 9250 on both storage & all client clusters.


Step 5 - Fetch and Merge Prometheus Configurations

Each bridge instance exposes a ready-made Prometheus scrape configuration at /prometheus.yml. Fetch both, then merge them into a single prometheus.yml using the script below.

5.1 Fetch each cluster's scrape config

The bridge exposes a ready-made Prometheus scrape config at /prometheus.yml. You can append a jobname_suffix query parameter to tag jobs with a cluster name or hostname, which helps distinguish targets in Prometheus:

# Storage cluster bridge
curl -k "https://<storage_cluster_ip>:9250/prometheus.yml?jobname_suffix=<clustername_or_hostname>" \
  -o /tmp/prom_storage.yml

# Client cluster bridge (also runs on port 9250)
curl -k "https://<client_cluster_ip>:9250/prometheus.yml?jobname_suffix=<clustername_or_hostname>" \
  -o /tmp/prom_client.yml

5.2 Run the merge script

Download merge_prometheus_configs.py from this wiki repository:

curl -O https://raw.githubusercontent.com/IBM/ibm-spectrum-scale-bridge-for-grafana/master/wiki/scripts/merge_prometheus_configs.py
chmod +x merge_prometheus_configs.py

Run it (all flags are optional - defaults match the standard layout):

python3 merge_prometheus_configs.py \
  --storage-yml          /tmp/prom_storage.yml \
  --client-yml           /tmp/prom_client.yml \
  --output               /etc/prometheus/prometheus.yml \
  --storage-cluster-name storage-cluster \
  --client-cluster-name  client-cluster

Validate and reload:

promtool check config /etc/prometheus/prometheus.yml && systemctl restart prometheus

Step 6 - Verify Endpoints and Metrics Collection

Bridge endpoints

# Storage bridge
curl -k https://<storage_cluster_ip>:9250/endpoints
curl -k https://<storage_cluster_ip>:9250/metadata/time

# Client bridge
curl -k https://<client_cluster_ip>:9250/endpoints
curl -s -k https://<client_cluster_ip>:9250/metadata/sensormetrics | python3 -m json.tool

Note: Open the Grafana UI at http://<storage_cluster_ip>:3000, add your Prometheus datasource pointing to http://<storage_cluster_ip>:9090 and start building dashboards to see all the storage and client metrics in a same dashboard. Use the clusterName label in PromQL to filter by cluster:

gpfs_health_status{clusterName="storage-cluster"}
gpfs_health_status{clusterName="client-cluster"}