Posts
Page 12 of 72
-
Node Exporter Internals That Actually Matter
Which 30 Node Exporter metrics matter for Prometheus monitoring. Skip the 200 noisy ones. Textfile collector trick included.
· Updated:10 min read -
Grafana OnCall + Webhooks: Paging Without PagerDuty
Self-host Grafana OnCall as a free PagerDuty alternative: escalation policies, Alertmanager webhooks, and phone alerts for your home lab (plus the 2026 catch).
· Updated:10 min read -
Healthchecks.io Self-Hosted: Cron Monitoring
Self-hosted Healthchecks.io for cron monitoring with dead-man-switch alerting, Docker setup, and backup integration. Never miss a silent failure again.
· Updated:9 min read -
Vector vs Fluent Bit: Modern Log Shippers
Vector vs Fluent Bit log shippers: Rust-based transforms and sinks vs lightweight C agent. Config, VRL, footprint, and when each wins.
· Updated:11 min read -
Push vs Pull Metrics: Pushgateway, Pushprox, and Why
Prometheus pull metrics are great until they aren't. When batch jobs, NAT, and ephemeral containers break scraping, Pushgateway, PushProx, and alternatives explained.
11 min read -
Blackbox Exporter: HTTP/TCP/DNS/ICMP Probes
Blackbox exporter Prometheus probes HTTP TCP DNS ICMP synthetic monitoring without paying for Pingdom or SaaS uptime tools
· Updated:9 min read -
Self-Host SigNoz: Install Guide
SigNoz killed install.sh and its bundled Compose files. Here's how to self-host it with Foundry, wire an app via OTLP, and set retention and alerts.
· Updated:18 min read -
SLO/SLI for Home Lab Services
Error budgets work on a home lab too. What an SLI and SLO actually are, picking an availability target you can hit, 28-day windows, and burn-rate alerts.
· Updated:10 min read -
Alertmanager Routing Trees That Don't Lie
Master Alertmanager routing trees: severity labeling, inhibition, grouping, and multi-receiver logic for reliable home lab alerts that work at 3 AM
9 min read -
Tempo + Grafana: Tracing for Self-Hosters
Distributed tracing with Tempo and Grafana for self-hosted multi-service stacks. OpenTelemetry, sampling, exemplars, and when tracing actually matters at home scale.
· Updated:9 min read -
Bots Ate 90% of My Worker Quota
A WordPress login bot burned 90% of my Cloudflare Workers free tier in two hours attacking a site that has never run PHP. Here's what actually stopped it.
13 min read -
Mimir + Grafana: Long-Term Prometheus Storage
Mimir for long-term Prometheus storage. S3-backed, horizontally scaled alternative to Thanos for capacity planning and compliance.
· Updated:8 min read