DEVOPS / A CONCEPT NOTE

Monitoring & Observability

knowing what your system is doing at all times

~90 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Monitoring collects metrics (CPU, latency, error rates) and triggers alerts when thresholds are breached. Observability goes further — with traces (distributed request paths), structured logs, and rich metrics — so you can ask arbitrary questions about system behavior without pre-defining every dashboard.

02 / FOLLOW THE MECHANISM

How a metric becomes an alert

  1. Exporter/node agent

    runs on every machine, collecting CPU, memory, disk, and custom app metrics and exposing them at /metrics.

  2. Prometheus server

    scrapes all targets at a configured interval (e.g., 15s) and stores metrics in a time-series DB.

  3. Alertmanager

    evaluates PromQL rules — if node_cpu_seconds_total > 0.9 for 5 minutes, fires an alert.

  4. Alertmanager

    deduplicates, groups, and routes the alert to PagerDuty, Slack, or email based on severity and team.

  5. On-call engineer

    receives the notification, opens Grafana, and investigates the dashboard or traces to find the root cause.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

validate Prometheus alerting rules

promtool check rules alerts.yml

EXAMPLE 02 · REFERENCE

query Prometheus API for target status

curl localhost:9090/api/v1/query?query=up

Explore command anatomy in the CLI lab