DEVOPS / A CONCEPT NOTE
Monitoring & Observability
knowing what your system is doing at all times
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Monitoring collects metrics (CPU, latency, error rates) and triggers alerts when thresholds are breached. Observability goes further — with traces (distributed request paths), structured logs, and rich metrics — so you can ask arbitrary questions about system behavior without pre-defining every dashboard.
02 / FOLLOW THE MECHANISM
How a metric becomes an alert
Exporter/node agent
runs on every machine, collecting CPU, memory, disk, and custom app metrics and exposing them at
/metrics.Prometheus server
scrapes all targets at a configured interval (e.g., 15s) and stores metrics in a time-series DB.
Alertmanager
evaluates PromQL rules — if
node_cpu_seconds_total > 0.9for 5 minutes, fires an alert.Alertmanager
deduplicates, groups, and routes the alert to PagerDuty, Slack, or email based on severity and team.
On-call engineer
receives the notification, opens Grafana, and investigates the dashboard or traces to find the root cause.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
validate Prometheus alerting rules
promtool check rules alerts.ymlquery Prometheus API for target status
curl localhost:9090/api/v1/query?query=up05 / CHECK YOURSELF
Could you explain Monitoring & Observability to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in Delivery & operationsConfiguration Managementautomating server setup and software configuration at scale