YOUR LEARNING PATH / 7 STAGES

Site Reliability Engineer

Keep systems fast, available, and resilient — with engineering, not heroics.

Learn · build · explain

The path, stage by stage.

7 stages

Open a stage for skills, a project, and references. Mark it complete when you can explain what you built and why.

  1. 01Systems, Linux & Networking Internals4 focus areas · 2 resources

    SREs debug the hardest problems, so deep systems knowledge is the foundation.

    What to learn

    • Linux internals: processes, memory, I/O, cgroups
    • Performance tools: top, htop, strace, perf
    • TCP/IP deep dive, DNS, TLS, HTTP/2 & gRPC
    • How load balancers, proxies, and caches work

    Put it into practice

    Your build

    Diagnose a slow Linux service using strace/perf and write up the root cause.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  2. 02Coding & Automation4 focus areas · 2 resources

    SRE = "what happens when a software engineer does ops." You must write real code.

    What to learn

    • Strong scripting + one real language (Go or Python)
    • Writing CLIs, automation, and small services
    • APIs, retries, timeouts, backoff & idempotency
    • Testing and code review discipline

    Put it into practice

    Your build

    Write a small Go/Python service with health checks, metrics, and graceful shutdown.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  3. 03Cloud, Containers & Kubernetes4 focus areas · 2 resources

    Modern reliability work happens on cloud-native, containerized platforms.

    What to learn

    • Docker + Kubernetes operations (not just deploys)
    • Cloud core services & managed Kubernetes
    • IaC with Terraform/OpenTofu
    • Capacity planning & autoscaling

    Put it into practice

    Your build

    Run an app on Kubernetes and intentionally kill pods to watch self-healing in action.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  4. 04Observability & the Three Pillars4 focus areas · 2 resources

    Reliability starts with measurement. Master metrics, logs, and traces.

    What to learn

    • Prometheus (PromQL) + Grafana dashboards
    • Logging pipelines (Loki / ELK)
    • Distributed tracing via OpenTelemetry
    • Designing actionable, low-noise alerts

    Put it into practice

    Your build

    Instrument your service with metrics + traces and build a RED/USE dashboard.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  5. 05SLIs, SLOs & Error Budgets4 focus areas · 2 resources

    The defining SRE skill: turning "is it reliable?" into measurable, agreed numbers.

    What to learn

    • Define SLIs that reflect user experience
    • Set SLOs and compute error budgets
    • Error-budget policies & balancing features vs reliability
    • Capacity & toil reduction targets

    Put it into practice

    Your build

    Pick an SLI for your service, set a 99.9% SLO, and chart the remaining error budget.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  6. 06Incident Response & Blameless Postmortems4 focus areas · 2 resources

    When things break, calm, structured response is what makes an SRE valuable.

    What to learn

    • On-call, paging, and incident command roles
    • Triage, mitigation-first thinking, comms
    • Blameless postmortems & action items
    • AI-assisted incident triage (AIOps) — emerging

    Put it into practice

    Your build

    Run a mock incident, then write a blameless postmortem with timeline + action items.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  7. 07Chaos Engineering & Resilience4 focus areas · 2 resources

    Don't wait for outages — inject failure on purpose and learn before customers do.

    What to learn

    • Failure modes: timeouts, retries, circuit breakers
    • Chaos experiments (Chaos Mesh / Litmus)
    • Load & soak testing, graceful degradation
    • DR strategy, backups & multi-region basics

    Put it into practice

    Your build

    Run a chaos experiment that injects latency, then add a timeout/retry to survive it.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

Extend the foundations

Ideas to connect as you go.

Return to these themes as the core skills become familiar.

Explore another direction.

DevOps EngineerBuild, automate, ship, and operate software reliably.Cloud EngineerDesign, deploy, and run workloads on the public cloud.AI / MLOps EngineerTake models — including LLMs — from notebook to reliable production.