Ravindra BagaleCourses & study guides

Chapter 13: Monitoring Basics

13.8 Alerting — one disk threshold that matters

Alert fatigue — so many noisy pages that humans mute everything. Prefer one clear alert you trust.

Worked idea: alert when disk used on / stays above a threshold (for example 85% or 90% — pick a lab number and document it) for several minutes, not for one noisy scrape.

Conceptually:

  1. Express "disk almost full" in PromQL (using filesystem metrics + mountpoint filter).
  2. Create an alert rule in Prometheus and/or Grafana (tooling differs by version — follow current docs).
  3. Route notification somewhere you control for labs (email you own, a private chat webhook you set up carefully — no real production secrets in Git).
  4. Trigger safely: create a large temporary file on a lab filesystem you own, watch the alert fire, then delete the file and confirm recovery.
# Sketch — fill then clean (ONLY on a disposable lab disk path you understand)
# Pick a directory with enough free space; do not fill a shared production volume.
dd if=/dev/zero of=/var/tmp/badge-api-fill.bin bs=1M count=512
df -h /var/tmp
# ... wait for scrape + alert ...
rm -f /var/tmp/badge-api-fill.bin
df -h /var/tmp

What you see: used space rises; alert goes firing; after rm, space returns and alert resolves. If nothing fires, check scrape interval, rule for: duration, and whether you filled the same mount the query watches.

Ravindra Bagale's Tip

Alert without runbook = anxiety. Notes madhe: "Disk high → check logs, docker system df, delete fill file, then investigate growth." Dhyan rakho!