PRACTICE TRACK / 10 QUESTIONS

Prometheus
Think it through.

Metrics collection, PromQL, alerting, and observability.

Choose a question, explain your approach, then reveal the supplied answer. Difficulty labels come from the existing question library.

10 questions

Answers stay closed until you choose to reveal them.

QUESTION 01PrometheusEasy

What is Prometheus and how does it collect metrics?

#
Reveal answer guidance

Prometheus is an open-source monitoring and alerting system. It collects metrics via a pull model — scraping HTTP endpoints (typically /metrics) on targets at configured intervals. Targets are discovered via service discovery (Kubernetes, Consul, EC2), static configs, or file-based SD. Each scrape stores samples with labels (key-value pairs) and a timestamp. Data is stored locally on disk with a custom time-series database using the TSDB format.

QUESTION 02PrometheusEasy

What metric types does Prometheus support?

#
Reveal answer guidance

Four core metric types: (1) Counter — cumulative value that only increases (total requests, errors). (2) Gauge — snapshot value that can go up or down (CPU, memory, queue size). (3) Histogram — samples observations in configurable buckets, also provides sum and count of observations (request latency). (4) Summary — similar to Histogram but calculates configurable quantiles on the client side. Each metric has a name and a set of labels identifying its dimensions.

QUESTION 03PrometheusMedium

What is the difference between rate() and increase() in PromQL?

#
Reveal answer guidance

Both calculate per-second rate of change for Counter metrics over a time range. rate(metric[5m]) returns the per-second average rate of increase. increase(metric[5m]) returns the total increase over the time range. Relationship: increase = rate * time_range_seconds. Both handle counter resets (restarts) correctly. Prefer rate() for alerting and dashboards because it produces consistent units (per second), and increase() for understanding total volume in a window.

QUESTION 04PrometheusMedium

How do relabel_configs differ from metric_relabel_configs?

#
Reveal answer guidance

relabel_configs runs before ingestion — modifies target labels (which target to scrape, how to identify it). Used for: filtering targets, adding/replacing instance labels, setting honor_labels. metric_relabel_configs runs after ingestion — modifies individual metric labels within scraped samples. Used for: dropping high-cardinality labels, renaming metrics, aggregating similar metrics. Misusing metric_relabel_configs can silently drop metrics — always test with a debug endpoint first.

QUESTION 05PrometheusMedium

What are recording rules and alerting rules in Prometheus?

#
Reveal answer guidance

Recording rules precompute frequently-used or expensive PromQL expressions and store them as new time series. Syntax: groups: [- name: node_rules, rules: [record: job:node_memory_util:avg, expr: avg(node_memory_MemTotal_bytes - node_memory_MemFree_bytes) / node_memory_MemTotal_bytes]]. Alerting rules define conditions that trigger alerts: alert: HighMemoryUsage, expr: job:node_memory_util:avg > 0.8, for: 5m, labels: {severity: warning}, annotations: {summary: "Memory usage above 80%"}. The for clause prevents flapping — alert fires only after sustained violation.

QUESTION 06PrometheusHard

How does Prometheus handle data storage, retention, and compaction?

#
Reveal answer guidance

Prometheus TSDB stores data in 2-hour blocks on disk under --storage.tsdb.path (default data/). Each block contains chunk files (compressed samples), an index (label → series mapping), and metadata. Retention: --storage.tsdb.retention.time=15d deletes blocks older than 15 days. Compaction merges smaller blocks into larger ones, deduplicates, and applies downsampling (5m resolution after 40 days, 1h after 10d with flags). Write-ahead log (WAL) in data/wal/ ensures durability — on crash, Prometheus replays the WAL. For high ingestion rates, tune --storage.tsdb.retention.size to limit total bytes instead of time.

QUESTION 07PrometheusHard

How does Thanos extend Prometheus for global querying and long-term storage?

#
Reveal answer guidance

Thanos adds: (1) Thanos Sidecar — ships Prometheus TSDB blocks to object storage (S3, GCS) and serves store API queries. (2) Thanos Store Gateway — serves historical data from object storage. (3) Thanos Query — provides a global Prometheus API endpoint that queries multiple Prometheus instances + Store Gateway simultaneously. (4) Thanos Compactor — deduplicates, downsampling, and retention enforcement on object storage. (5) Thanos Ruler — evaluates recording/alerting rules without scraping. Architecture: each Prometheus runs a sidecar; Thanos Query federates across all of them. Deduplication via --query.replica-label flag handles HA Prometheus pairs.

QUESTION 08PrometheusHard

How do you handle high cardinality in Prometheus and prevent it from causing performance issues?

#
Reveal answer guidance

High cardinality (millions of unique label combinations) causes high memory usage, slow queries, and OOM kills. Prevention: (1) Use metric_relabel_configs to drop high-churn labels (e.g., pod IDs, request IDs, user IDs). (2) Aggregate before storing — pre-compute with recording rules. (3) Enable --storage.tsdb.max-exemplars carefully. (4) Use --storage.tsdb.sample-count-per-block limits. (5) Monitor cardinality with prometheus_tsdb_series_count. (6) Set --storage.tsdb.retention.time strictly. (7) Use promtool tsdb analyze to identify top label pair cardinality. Remediation: drop offending labels via relabeling or aggregate unwanted series. For user-facing label dimensions, use a separate Prometheus with different retention.

QUESTION 09PrometheusMedium

How does Prometheus service discovery work in Kubernetes?

#
Reveal answer guidance

The Prometheus Operator or kubernetes_sd_configs discovers targets via the Kubernetes API. Four roles: (1) node — discovers all nodes, grabs node-exporter metrics. (2) service — discovers Services, useful for Service-level metrics. (3) pod — discovers all pods, anno-based scrape config. (4) endpoints — discovers Endpoints (the default for kubelet). Relabeling filters targets: __meta_kubernetes_pod_annotation_prometheus_io_scrape = "true". The __metrics_path__ label sets the scrape path. __meta_kubernetes_namespace identifies the namespace.

QUESTION 10PrometheusMedium

How does the Prometheus Operator simplify monitoring in Kubernetes?

#
Reveal answer guidance

The Prometheus Operator introduces CRDs: Prometheus (deploy/manage Prometheus instances), ServiceMonitor (define scrape targets declaratively), PodMonitor (scrape pods directly), PrometheusRule (alerting and recording rules), Alertmanager (manage Alertmanager instances). Benefits: (1) Auto-discovers targets via label selectors. (2) Manages Prometheus upgrades, scaling, and configuration. (3) Provides kube-prometheus-stack (kube-prometheus) — a complete monitoring stack with Grafana dashboards, node-exporter, kube-state-metrics, and default alerting rules.

CONTINUE PRACTICING

Try another perspective.