SLO/SLI for Home Lab Services
SRE error budgets for home lab. Build Grafana SLO dashboards, understand SLI/SLO, pick sane availability targets.
All the articles with the tag "observability".
SRE error budgets for home lab. Build Grafana SLO dashboards, understand SLI/SLO, pick sane availability targets.
Master Alertmanager routing trees: severity labeling, inhibition, grouping, and multi-receiver logic for reliable home lab alerts that work at 3 AM
Distributed tracing with Tempo and Grafana for self-hosted multi-service stacks. OpenTelemetry, sampling, exemplars, and when tracing actually matters at home scale.
Mimir for long-term Prometheus storage. S3-backed, horizontally scaled alternative to Thanos for capacity planning and compliance.
Promtail vs Vector: which log shipping agent fits your setup? Loki's purpose-built simplicity or Vector's transform power and multiple sinks.
Prometheus federation for multi-site home labs: hierarchical scraping, cross-cluster aggregation, federation vs remote_write, practical config.
Jaeger does distributed tracing only; SigNoz bundles traces, metrics, and logs. Which should you self-host? Here's the honest breakdown.
SigNoz gives you logs, metrics, and traces in one ClickHouse-backed app. The Grafana LGTM stack gives you four. Here's how to pick the right one.
ClickHouse can store TB of logs on one node, query in seconds, and outscale Loki and Elasticsearch in raw cost. Here's how to wire it up with Vector and Grafana.
When Prometheus is overkill: a push-based metrics stack for IoT, smart home, and homelab data. Telegraf, InfluxDB, and Grafana, working together end-to-end.
Two OpenTelemetry-first observability platforms backed by ClickHouse. Here's how SigNoz and Uptrace compare for self-hosted homelab and small-team use.
Promtail is deprecated. Here's how to migrate your scrape configs, pipeline stages, and Docker log setups to Grafana Alloy without losing your mind or data.