Availability
Loss of an SNMP scrape is detected independently of the visualization layer.
A legacy LenovoEMC appliance and TrueNAS CORE brought into one operational model using SNMPv3, vendor MIBs, Prometheus recording rules, alerting and a compact Grafana storage dashboard.
The same environment contains two very different NAS platforms: a closed legacy LenovoEMC appliance and a TrueNAS CORE system built around ZFS. Their OIDs, MIB structure and telemetry capabilities differ substantially.
The objective was not to collect a handful of SNMP values, but to create one operational view covering availability, capacity, RAID/ZFS state, temperatures, IOPS, throughput, network interfaces and actionable alerts.
Each device is scraped with its own SNMP modules. Once metrics pass through snmp_exporter, Prometheus applies the same operational patterns to both platforms.
A dedicated TrueNAS module based on FREENAS-MIB exposes ZFS pool size and state, read/write counters, ARC data and disk temperatures. IF-MIB was later added for interface state, traffic counters, errors and discards.
LenovoEMC capacity comes from HOST-RESOURCES-MIB, while RAID, physical disk state, temperatures and fan speeds come from the vendor tree. A full SNMP walk is not exported into Prometheus; the module stays intentionally narrow to reduce noise and cardinality.
- targets:
- 10.20.10.21
labels:
device: nas-truenas-01
vendor: ixsystems
role: storage
snmp_auth: truenas_v3
snmp_module: system,if_mib,truenas_corerate(truenas_zpool_read_bytes_total[5m])
rate(truenas_zpool_write_bytes_total[5m])
rate(ifHCInOctets{ifName="storage0"}[5m])
rate(ifHCOutOctets{ifName="storage0"}[5m])ONLINE/DEGRADED/FAULTED for ZFS plus vendor-specific RAID and physical disk state for the legacy appliance.
Used, available, total and percentage values with allocation units normalized to bytes.
Read/write throughput, IOPS, ARC, temperatures and network RX/TX for active interfaces.
FREENAS-MIB represents pool capacity through allocation units. Conversion to bytes and percentages is handled by Prometheus recording rules. Grafana and alerts therefore consume normalized metrics instead of repeating vendor-specific expressions.
This layer makes future changes easier: the collection or recording rule can change while dashboards and operational alerting remain comparatively stable.
- record: storage:truenas:zpool_used_percent
expr: |
100 * truenas_zpool_used_units
/ truenas_zpool_size_units- alert: TrueNASZpoolNotOnline
expr: truenas_zpool_health != 0
for: 2m
labels:
severity: critical
category: storageLoss of an SNMP scrape is detected independently of the visualization layer.
Any ZFS pool state other than ONLINE and abnormal RAID or disk state require attention.
80%, 90% and 95% thresholds separate warning, critical and almost-full conditions.
HDD and NVMe use different operational thresholds to avoid noisy alerts.
The first dashboard revision contained many historical graphs. It was technically detailed but difficult to scan as a NOC view. The final design was rebuilt around current operational state.
One screen now covers availability, alerts, temperatures, ARC hit ratio, ZFS/RAID health, capacity, read/write throughput, IOPS, interface state, network RX/TX and errors/discards. Detailed history can remain in a separate drill-down dashboard.
Availability · alerts · temperatures · ARC · interfaces.
Pool capacity · ZFS health · RAID · physical disks.
Read/write · IOPS · network RX/TX · errors/discards.
A legacy appliance and TrueNAS now share one monitoring pipeline without agents on the storage systems. Vendor-specific SNMP stays at collection time; normalized metrics, common alerting patterns and a compact Grafana overview sit above it.
← Back to projects