You Imported Dashboard 1860 and Now You Scroll for a Living
It started innocently. You installed node_exporter, imported Node Exporter Full (dashboard ID 1860), and felt like a real operator. Then you added a few panels for the new container host. Then someone on a forum pasted a “better CPU graph.” Now your home dashboard is a vertical wall of charts, and when the NAS goes sideways at 2 AM you scroll past forty graphs to find out which one is red.
The fix is a rule, not a better layout. Every panel on the main dashboard must either answer a question you ask out loud (“is the host healthy?”) or feed an alert. Everything else moves to a drill-down dashboard you open on purpose. To decide which ten survive, borrow two methods that already did the thinking for you: USE for hosts and RED for services.
The current Grafana release is v13.2.3 as of this writing, but nothing here depends on a recent version. Dashboards, collapsed rows, data links, and file provisioning have worked this way for years.
USE and RED in Sixty Seconds
Brendan Gregg’s USE method is for resources (CPU, memory, disks, NICs). For each resource, check three things:
- Utilization: how busy the resource is.
- Saturation: how much work is queued waiting for it.
- Errors: how many things went wrong.
Tom Wilkie’s RED method is for request-driven services. For each service, track:
- Rate: requests per second.
- Errors: how many of those requests fail.
- Duration: how long they take.
Hosts get USE, services get RED. Anything that fits neither is probably a curiosity, and curiosities belong on a drill-down. A 4,000 RPM reading on a fan sensor is fascinating. It is also not a question you ask at 2 AM.
Step 1: Audit What You Actually Have
Before you delete anything, find out what is in the dashboard. Grafana Enterprise and Grafana Cloud include dashboard usage insights that show views per dashboard. Grafana OSS does not have that feature, so you improvise. There are two cheap ways.
First, count dashboard hits in your reverse proxy log. This gives you per-dashboard views, not per-panel, but it tells you which dashboards are dead weight. This assumes a combined-format access log:
awk '{print $7}' /var/log/proxy/access.log \ | grep '^/d/' | cut -d/ -f3 | sort | uniq -c | sort -rnSecond, dump every panel title and query through the HTTP API. Create a service account with a Viewer role, grab its token, and run this:
G=http://grafana.lan:3000UID_=rYdddlPWk # the uid from the dashboard URL
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \ "$G/api/dashboards/uid/$UID_" \ | jq -r '[.. | objects | select(has("targets") and has("title"))] | .[] | [.title, (.targets[0].expr // "-")] | @tsv' \ > panels.tsv
wc -l panels.tsvThe .. recursion matters because collapsed rows hide their panels inside a nested panels array. Without it, you undercount. Now find the redundancy:
grep -o 'node_[a-z_]*' panels.tsv | sort | uniq -c | sort -rn | headIf node_cpu_seconds_total shows up in fourteen panels, you do not have fourteen insights. You have one insight drawn fourteen ways: stacked, per-core, per-mode, as a gauge, as a stat, as a bar. Pick the one that answers “is the CPU the problem?” and mark the rest for the drill-down pile.
The audit gets you a candidate list. The human part is walking each panel and asking one question.
Step 2: The Rule That Does the Pruning
For each panel, ask: what decision does this change? Three valid answers:
- It answers a question you ask on a bad day (“is it the disk or the network?”).
- It backs an alert, so when the alert fires, you land on the panel that shows why.
- It is the one place a hard number lives (like free disk space) that you check weekly.
“It looks cool” and “the original author included it” are not answers. A dashboard is a cockpit, not a museum. Moving a panel off the main dashboard does not delete the data. Prometheus still has every series, and Explore can query it on demand.
This is the forklift-and-couch problem in reverse. You do not need a forklift of 50 panels to check whether the couch is where you left it.
Step 3: The Ten Panels
Seven for the host, three for one service. Create two dashboard variables first: instance (query label_values(node_uname_info, instance)) and, for the service, a job variable or a hardcoded job name. The Grafana variables post in this blog covers the mechanics.
Host: USE (7 panels)
1. CPU utilization (%). The standard idle-subtraction:
100 * (1 - avg by (instance) ( rate(node_cpu_seconds_total{mode="idle", instance=~"$instance"}[5m])))2. CPU saturation (pressure). Load average is the classic saturation signal, but it conflates CPU and disk waits on Linux. Pressure Stall Information (PSI) is cleaner. The node_pressure_cpu_waiting_seconds_total counter counts seconds that tasks waited for a CPU. A rate of 1 means a task was waiting the whole second, so multiply by 100 for a percentage of time stalled:
100 * rate(node_pressure_cpu_waiting_seconds_total{instance=~"$instance"}[5m])PSI needs Linux kernel 4.20 or newer with CONFIG_PSI enabled (some distro kernels also want psi=1 on the kernel command line), and node_exporter’s pressure collector (on by default). If the panel is empty, check /proc/pressure/cpu exists. If it does not, fall back to node_load1 / count by (instance) (node_cpu_seconds_total{mode="idle"}), which is load per core. Above 1 sustained means a queue.
3. Memory utilization (%). Use MemAvailable, not MemFree. Linux fills free memory with page cache on purpose, so “free” is always low and always misleading:
100 * (1 - node_memory_MemAvailable_bytes{instance=~"$instance"} / node_memory_MemTotal_bytes{instance=~"$instance"})4. Memory saturation (pressure). Same PSI idea, memory flavor. Sustained nonzero values mean the box is reclaiming or swapping hard enough that tasks stall:
100 * rate(node_pressure_memory_waiting_seconds_total{instance=~"$instance"}[5m])5. Filesystem fullness (%). Use avail, not free. avail excludes blocks reserved for root, so it matches what df shows and what your services actually get. Filter out the pseudo-filesystems:
max by (instance, mountpoint) ( 100 * (1 - node_filesystem_avail_bytes{instance=~"$instance", fstype!~"tmpfs|overlay|squashfs|ramfs"} / node_filesystem_size_bytes{instance=~"$instance", fstype!~"tmpfs|overlay|squashfs|ramfs"}))6. Disk saturation (I/O pressure). Disk utilization (rate(node_disk_io_time_seconds_total[5m]), the fraction of time the device had I/O in flight) is easy to misread on SSDs and NVMe, which service many requests in parallel. I/O pressure tells you whether tasks actually waited:
100 * rate(node_pressure_io_waiting_seconds_total{instance=~"$instance"}[5m])7. Network errors and drops. One panel, summed, because you only care whether it is nonzero. Exclude loopback:
sum by (instance) ( rate(node_network_receive_errs_total{instance=~"$instance", device!="lo"}[5m]) + rate(node_network_transmit_errs_total{instance=~"$instance", device!="lo"}[5m]) + rate(node_network_receive_drop_total{instance=~"$instance", device!="lo"}[5m]) + rate(node_network_transmit_drop_total{instance=~"$instance", device!="lo"}[5m]))You will notice USE has cells this layout skips: CPU errors, memory errors, per-NIC utilization. On a homelab, CPU and ECC memory errors show up in dmesg, not in a graph you check daily, and NIC throughput is a drill-down question. USE is a checklist to consider, not a quota to fill.
Service: RED (3 panels)
The queries below assume your service exports a standard Prometheus histogram named http_request_duration_seconds, which gives you _bucket, _count, and _sum series. Your service will use whatever name and labels its client library chose. Swap the metric name, job, and the status label to match what yours exports. Check with curl localhost:PORT/metrics | grep duration.
8. Rate. The _count series increments once per request, so its rate is requests per second:
sum(rate(http_request_duration_seconds_count{job="myapp"}[5m]))9. Errors. Ratio of 5xx to all requests. When traffic is zero, this divides zero by zero and shows no data, which is fine. Do not “fix” it with or vector(0) unless you want a flat zero that hides a dead service:
sum(rate(http_request_duration_seconds_count{job="myapp", status=~"5.."}[5m]))/sum(rate(http_request_duration_seconds_count{job="myapp"}[5m]))10. Duration. The 95th percentile from the buckets. Always aggregate with sum by (le) before histogram_quantile:
histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket{job="myapp"}[5m])))Ten panels, three rows of thought: is the host busy, is the host stalled, is the service healthy. You can read that in the time it takes to sip coffee.
Step 4: Layout That Reads Like a Sentence
Arrange the panels so the eye does the work:
- Top row, host: CPU utilization and saturation side by side, then memory utilization and saturation. Utilization left, saturation right, always. Your brain learns the pattern in a day.
- Second row, host: filesystem fullness, disk I/O pressure, network errors.
- Third row, service: rate, errors, duration, left to right in the order the acronym says.
Set thresholds so color does the reading. On the pressure panels, go yellow around 10% and red around 25%. Those are starting points, not gospel; tune them against a week of your own data. A panel that is green when healthy and red when not does not need a legend.
Step 5: Move the Rest to Drill-Downs
The other forty panels are not garbage. They are the second-level questions: “CPU is busy, but which mode?” Keep your imported 1860 dashboard exactly as it is and treat it as the drill-down. Link into it from the ten.
Add a data link on a panel. In the panel JSON, it lives under fieldConfig.defaults.links:
"fieldConfig": { "defaults": { "links": [ { "title": "CPU detail", "url": "/d/rYdddlPWk/node-exporter-full?var-job=${__field.labels.job}&var-node=${__field.labels.instance}&${__url_time_range}", "targetBlank": false } ] }}${__field.labels.instance} pulls the clicked series’ instance label value. ${__url_time_range} carries your current time range into the target dashboard, so you land on the same window you were staring at. Dashboard 1860 chains its dropdowns (job, then nodename, then node), so pass var-job too, or the destination can reset node to its default. If it still lands on the wrong host, set the nodename dropdown once. Every var- name must match a variable on the destination dashboard. Open it, change a dropdown, and read the names out of the URL.
Not every panel needs a link. For a few, a collapsed row is simpler: group the second-tier panels inside the same dashboard under a row that starts closed. In JSON, a collapsed row carries its children inside it:
{ "type": "row", "title": "Details (click to expand)", "collapsed": true, "gridPos": { "h": 1, "w": 24, "x": 0, "y": 20 }, "panels": [ { "type": "timeseries", "title": "CPU by mode", "gridPos": { "h": 8, "w": 12, "x": 0, "y": 21 } } ]}Collapsed rows have a cost: Grafana does not run queries for panels inside a collapsed row, so they cost nothing until opened. Use them for “same dashboard, one more click.” Use data links for “different dashboard, different question.” Either way, the front page stays at ten.
Step 6: Keep It as Code
A pruned dashboard that anyone can edit through the UI will regrow within a month. Put the JSON in git and let Grafana load it from disk.
Export the dashboard JSON with the API, drop the numeric id, and commit it:
curl -s -H "Authorization: Bearer $GRAFANA_TOKEN" \ "$G/api/dashboards/uid/$UID_" \ | jq '.dashboard | del(.id)' > dashboards/homelab-overview.json
git add dashboards/homelab-overview.jsongit commit -m "dashboards: trim overview to 10 panels (USE host, RED myapp)"Then tell Grafana to watch the directory with a provisioning file:
apiVersion: 1providers: - name: homelab orgId: 1 folder: Homelab type: file disableDeletion: true allowUiUpdates: false updateIntervalSeconds: 30 options: path: /var/lib/grafana/dashboardsallowUiUpdates: false means UI edits to this dashboard cannot be saved. Grafana shows a “Cannot save provisioned dashboard” dialog when you try, which is the point. You edit the JSON, commit, and within updateIntervalSeconds Grafana picks up the change. Mount both paths in Compose:
services: grafana: image: grafana/grafana-oss:latest volumes: - ./provisioning/dashboards:/etc/grafana/provisioning/dashboards:ro - ./dashboards:/var/lib/grafana/dashboards:roThe practical workflow: prototype in the UI on a scratch copy, export JSON, commit, and let the provisioned version be the real one. Your 2 AM self will thank you when a bad edit is a git revert away rather than a mystery.
Step 7: Make Every Panel Earn Its Spot (Alerts)
The rule says a panel either answers a question or backs an alert. Close the loop on the second half. Pick the panels where a number crossing a line means “wake me up” and write the alert against the same query:
groups: - name: host-use rules: - alert: FilesystemFillingUp expr: predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay|squashfs|ramfs"}[6h], 24 * 3600) < 0 for: 30m labels: severity: warning annotations: summary: "{{ $labels.mountpoint }} on {{ $labels.instance }} fills within 24h" - alert: MemoryPressureSustained expr: rate(node_pressure_memory_waiting_seconds_total[5m]) > 0.25 for: 10m labels: severity: warningThe memory alert fires when tasks wait on memory more than 25% of the time for ten minutes, the same yellow-to-red line your panel uses. Now when the page arrives, the panel on your dashboard shows the same thing the alert saw. If a panel has no question and no alert after the audit, it moves to the drill-down. No appeals.
Common Questions
How many panels should a Grafana dashboard have?
A Grafana dashboard should hold about ten panels on its first screen, the count that fits one monitor without scrolling. Past that, you stop reading each panel and start scanning for red. Ten panels also map cleanly to USE for one host plus RED for one service.
Does Grafana OSS have dashboard usage insights?
No. Dashboard usage insights, which show views per dashboard, ship with Grafana Enterprise and Grafana Cloud, not with Grafana OSS. On OSS, count /d/ requests in your reverse proxy access log for per-dashboard views. Per-panel view counts are not available in OSS.
Why is my node_pressure metric missing?
The node_pressure_* metrics are missing when the kernel lacks pressure stall information. PSI needs Linux 4.20 or newer with CONFIG_PSI enabled, and the file /proc/pressure/cpu must exist. Some kernels also need psi=1 on the boot command line. node_exporter’s pressure collector is on by default. Some container or older kernel setups simply do not expose PSI.
Can I still edit a provisioned Grafana dashboard in the UI?
Only if the provider sets allowUiUpdates: true. With the default of false, Grafana refuses to save UI edits to a file-provisioned dashboard. You edit the JSON in git instead, and Grafana reloads it after updateIntervalSeconds. Export from a scratch copy if you prefer drawing panels with a mouse.