DEVOPS / A CONCEPT NOTE

Incident Management

the structured process of detecting, responding to, and learning from outages

~85 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

When a production incident happens — site down, errors spiking, data corrupted — you need a process, not panic. Incident management means: declare severity, assemble a response team (commander, comms, responders), communicate in a dedicated channel, mitigate (rollback, scale up, feature flag off), and follow up with a postmortem. Blameless culture is non-negotiable.

02 / FOLLOW THE MECHANISM

How an incident unfolds

  1. Alert

    PagerDuty/OpsGenie pages the on-call engineer because an SLO burn rate alert is firing.

  2. Triage

    the engineer acknowledges, assesses severity (SEV1 = customer-impacting outage, SEV2 = degraded, SEV3 = minor).

  3. Response

    if SEV1/SEV2, the incident commander is assigned, a war room (Slack/Discord) is created, and responders jump in.

  4. Mitigation

    priority is to restore service — rollback the deploy, scale out, enable a feature flag, or run a mitigation script.

  5. Postmortem

    within 48h, a blameless root cause analysis is written with action items to prevent recurrence.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

acknowledge a PagerDuty incident

pd ack INCIDENT_ID

EXAMPLE 02 · REFERENCE

resolve an incident via API

curl -X POST -d '{"status":"resolved"}' https://api.pagerduty.com/incidents/INCIDENT_ID

Explore command anatomy in the CLI lab