YOUR LEARNING PATH / 7 STAGES

AI / MLOps Engineer

Take models — including LLMs — from notebook to reliable production.

Learn · build · explain

The path, stage by stage.

7 stages

Open a stage for skills, a project, and references. Mark it complete when you can explain what you built and why.

  1. 01Python, Data & Engineering Foundations4 focus areas · 2 resources

    MLOps is software engineering applied to ML. Strong Python + data skills come first.

    What to learn

    • Python: typing, packaging, virtualenvs, testing
    • Git, Docker, and the CLI (see the DevOps track)
    • NumPy, Pandas, and data wrangling
    • SQL and basic data pipelines

    Put it into practice

    Your build

    Build a clean Python package with tests that loads, cleans, and summarizes a public dataset.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  2. 02Machine Learning Fundamentals4 focus areas · 2 resources

    Understand how models train & fail so you can operate them, even if you don't build them.

    What to learn

    • Supervised vs unsupervised learning, train/val/test
    • Overfitting, metrics (accuracy, precision/recall, AUC)
    • scikit-learn end-to-end workflow
    • Intro to neural nets (PyTorch basics)

    Put it into practice

    Your build

    Train, evaluate, and save a scikit-learn model, then load it in a separate script to predict.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  3. 03LLMs, RAG & Prompt/Context Engineering5 focus areas · 2 resources

    The hottest part of the field. Most AI jobs today are about applying LLMs well.

    What to learn

    • How LLMs work: tokens, context windows, temperature
    • Prompt & context engineering; structured outputs / tool use
    • Retrieval-Augmented Generation (RAG) end-to-end
    • Vector databases & embeddings (pgvector, Qdrant)
    • Agentic patterns & function/tool calling

    Put it into practice

    Your build

    Build a RAG chatbot over your own docs using an LLM API + a vector store.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  4. 04Model Serving & Inference4 focus areas · 2 resources

    Turning a model into a fast, scalable API is the core MLOps deliverable.

    What to learn

    • Wrap models in an API (FastAPI / BentoML)
    • Containerize and deploy inference services
    • High-throughput LLM serving with vLLM / TGI
    • Batching, quantization, and GPU vs CPU trade-offs

    Put it into practice

    Your build

    Serve an open model with vLLM in a container and load-test it with a simple benchmark.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  5. 05ML Pipelines, Tracking & Registries4 focus areas · 2 resources

    Reproducibility is what separates a demo from a production ML system.

    What to learn

    • Experiment tracking (MLflow / Weights & Biases)
    • Data & model versioning (DVC, model registry)
    • Pipeline orchestration (Airflow / Prefect / Kubeflow)
    • Feature stores (concept) & dataset lineage

    Put it into practice

    Your build

    Track 3 training runs in MLflow and register the best model in its model registry.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  6. 06Cloud, Containers & GPU Infrastructure4 focus areas · 2 resources

    Models run on real infra. Know enough cloud + Kubernetes to deploy and scale them.

    What to learn

    • Docker + Kubernetes for ML workloads
    • Cloud GPU instances & spot/preemptible cost control
    • Managed AI platforms (Bedrock, Vertex AI, SageMaker)
    • CI/CD for ML (CML, GitHub Actions for training/deploy)

    Put it into practice

    Your build

    Deploy your inference service to a managed container platform with autoscaling.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

  7. 07Evaluation, Monitoring & Guardrails4 focus areas · 2 resources

    LLMs are non-deterministic — evals and monitoring are the new unit tests.

    What to learn

    • Offline & online evals (LLM-as-judge, golden datasets)
    • Drift detection & data/model monitoring
    • Tracing LLM apps (LangSmith / OpenTelemetry GenAI)
    • Safety: guardrails, PII handling, prompt-injection defense

    Put it into practice

    Your build

    Add an eval suite + tracing to your RAG app and catch one regression with it.

    Keep notes on the setup, result, and one trade-off you made.

    Resources & documentation

Extend the foundations

Ideas to connect as you go.

Return to these themes as the core skills become familiar.

Explore another direction.

DevOps EngineerBuild, automate, ship, and operate software reliably.Cloud EngineerDesign, deploy, and run workloads on the public cloud.Site Reliability EngineerKeep systems fast, available, and resilient — with engineering, not heroics.