YOUR LEARNING PATH / 7 STAGES
Site Reliability Engineer
Keep systems fast, available, and resilient — with engineering, not heroics.
Learn · build · explain
The path, stage by stage.
Open a stage for skills, a project, and references. Mark it complete when you can explain what you built and why.
01Systems, Linux & Networking Internals4 focus areas · 2 resources
SREs debug the hardest problems, so deep systems knowledge is the foundation.
What to learn
- Linux internals: processes, memory, I/O, cgroups
- Performance tools: top, htop, strace, perf
- TCP/IP deep dive, DNS, TLS, HTTP/2 & gRPC
- How load balancers, proxies, and caches work
Put it into practice
Your build
Diagnose a slow Linux service using strace/perf and write up the root cause.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
02Coding & Automation4 focus areas · 2 resources
SRE = "what happens when a software engineer does ops." You must write real code.
What to learn
- Strong scripting + one real language (Go or Python)
- Writing CLIs, automation, and small services
- APIs, retries, timeouts, backoff & idempotency
- Testing and code review discipline
Put it into practice
Your build
Write a small Go/Python service with health checks, metrics, and graceful shutdown.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
03Cloud, Containers & Kubernetes4 focus areas · 2 resources
Modern reliability work happens on cloud-native, containerized platforms.
What to learn
- Docker + Kubernetes operations (not just deploys)
- Cloud core services & managed Kubernetes
- IaC with Terraform/OpenTofu
- Capacity planning & autoscaling
Put it into practice
Your build
Run an app on Kubernetes and intentionally kill pods to watch self-healing in action.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
04Observability & the Three Pillars4 focus areas · 2 resources
Reliability starts with measurement. Master metrics, logs, and traces.
What to learn
- Prometheus (PromQL) + Grafana dashboards
- Logging pipelines (Loki / ELK)
- Distributed tracing via OpenTelemetry
- Designing actionable, low-noise alerts
Put it into practice
Your build
Instrument your service with metrics + traces and build a RED/USE dashboard.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
05SLIs, SLOs & Error Budgets4 focus areas · 2 resources
The defining SRE skill: turning "is it reliable?" into measurable, agreed numbers.
What to learn
- Define SLIs that reflect user experience
- Set SLOs and compute error budgets
- Error-budget policies & balancing features vs reliability
- Capacity & toil reduction targets
Put it into practice
Your build
Pick an SLI for your service, set a 99.9% SLO, and chart the remaining error budget.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
06Incident Response & Blameless Postmortems4 focus areas · 2 resources
When things break, calm, structured response is what makes an SRE valuable.
What to learn
- On-call, paging, and incident command roles
- Triage, mitigation-first thinking, comms
- Blameless postmortems & action items
- AI-assisted incident triage (AIOps) — emerging
Put it into practice
Your build
Run a mock incident, then write a blameless postmortem with timeline + action items.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
07Chaos Engineering & Resilience4 focus areas · 2 resources
Don't wait for outages — inject failure on purpose and learn before customers do.
What to learn
- Failure modes: timeouts, retries, circuit breakers
- Chaos experiments (Chaos Mesh / Litmus)
- Load & soak testing, graceful degradation
- DR strategy, backups & multi-region basics
Put it into practice
Your build
Run a chaos experiment that injects latency, then add a timeout/retry to survive it.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
Extend the foundations
Ideas to connect as you go.
Return to these themes as the core skills become familiar.
- SLO-driven engineering & error budgets
- OpenTelemetry as the observability standard
- AIOps & AI-assisted incident response
- Chaos engineering & resilience testing