DEVOPS / A CONCEPT NOTE
Incident Management
the structured process of detecting, responding to, and learning from outages
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
When a production incident happens — site down, errors spiking, data corrupted — you need a process, not panic. Incident management means: declare severity, assemble a response team (commander, comms, responders), communicate in a dedicated channel, mitigate (rollback, scale up, feature flag off), and follow up with a postmortem. Blameless culture is non-negotiable.
02 / FOLLOW THE MECHANISM
How an incident unfolds
Alert
PagerDuty/OpsGenie pages the on-call engineer because an SLO burn rate alert is firing.
Triage
the engineer acknowledges, assesses severity (SEV1 = customer-impacting outage, SEV2 = degraded, SEV3 = minor).
Response
if SEV1/SEV2, the incident commander is assigned, a war room (Slack/Discord) is created, and responders jump in.
Mitigation
priority is to restore service — rollback the deploy, scale out, enable a feature flag, or run a mitigation script.
Postmortem
within 48h, a blameless root cause analysis is written with action items to prevent recurrence.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
acknowledge a PagerDuty incident
pd ack INCIDENT_IDresolve an incident via API
curl -X POST -d '{"status":"resolved"}' https://api.pagerduty.com/incidents/INCIDENT_ID05 / CHECK YOURSELF
Could you explain Incident Management to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in Delivery & operationsService Meshmanaging microservice traffic, security, and observability out-of-band