13.13 Lab — disk metric path and a triggered alert
Lab
Build the primary monitoring path for a badge API lab host and prove one disk alert.
- Run Node Exporter; confirm
/metricsresponds on your chosen port. - Run Prometheus with a scrape job for that exporter; confirm target UP.
- Run Grafana; add Prometheus data source; create a small dashboard (CPU, memory, disk %).
- Define one disk-used alert with a sustained duration.
- Trigger safely with a temporary fill file on a lab path; confirm alert fires.
- Delete the fill file; confirm alert resolves and disk recovers.
- Write a five-line runbook in your notes (not committed secrets).
- Optional: skim CloudWatch docs and write three bullets on how the same disk story would look there — no need to implement both.
Practice task
Explain golden signals to a project manager in six lines, mapping latency/traffic/errors/saturation to badge API peak traffic and a full disk.
Thodkyaat sangaycha tar
- Monitor so you see trouble before customers do — bike dashboard before the highway stop.
- Golden signals: latency, traffic, errors, saturation; host metrics explain saturation.
- Primary path: Node Exporter → Prometheus → Grafana; CloudWatch is the AWS-native alternate (link AWS course for depth).
- Prefer few clear panels and one trustworthy alert over noisy walls.
- Trigger disk alerts safely on lab disks and always clean up; ports are published ports — do not expose admin UIs publicly without care.
- No invented AWS prices or free-tier claims — stop idle labs.
Samajla ka? Nasel tar lab (13.13) scrape UP → dashboard → disk alert fire/resolve paryant sodu naka. Aata pudhe jaauya logging aani light observability kade.