YOUR LEARNING PATH / 7 STAGES
AI / MLOps Engineer
Take models — including LLMs — from notebook to reliable production.
Learn · build · explain
The path, stage by stage.
Open a stage for skills, a project, and references. Mark it complete when you can explain what you built and why.
01Python, Data & Engineering Foundations4 focus areas · 2 resources
MLOps is software engineering applied to ML. Strong Python + data skills come first.
What to learn
- Python: typing, packaging, virtualenvs, testing
- Git, Docker, and the CLI (see the DevOps track)
- NumPy, Pandas, and data wrangling
- SQL and basic data pipelines
Put it into practice
Your build
Build a clean Python package with tests that loads, cleans, and summarizes a public dataset.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
02Machine Learning Fundamentals4 focus areas · 2 resources
Understand how models train & fail so you can operate them, even if you don't build them.
What to learn
- Supervised vs unsupervised learning, train/val/test
- Overfitting, metrics (accuracy, precision/recall, AUC)
- scikit-learn end-to-end workflow
- Intro to neural nets (PyTorch basics)
Put it into practice
Your build
Train, evaluate, and save a scikit-learn model, then load it in a separate script to predict.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
03LLMs, RAG & Prompt/Context Engineering5 focus areas · 2 resources
The hottest part of the field. Most AI jobs today are about applying LLMs well.
What to learn
- How LLMs work: tokens, context windows, temperature
- Prompt & context engineering; structured outputs / tool use
- Retrieval-Augmented Generation (RAG) end-to-end
- Vector databases & embeddings (pgvector, Qdrant)
- Agentic patterns & function/tool calling
Put it into practice
Your build
Build a RAG chatbot over your own docs using an LLM API + a vector store.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
04Model Serving & Inference4 focus areas · 2 resources
Turning a model into a fast, scalable API is the core MLOps deliverable.
What to learn
- Wrap models in an API (FastAPI / BentoML)
- Containerize and deploy inference services
- High-throughput LLM serving with vLLM / TGI
- Batching, quantization, and GPU vs CPU trade-offs
Put it into practice
Your build
Serve an open model with vLLM in a container and load-test it with a simple benchmark.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
05ML Pipelines, Tracking & Registries4 focus areas · 2 resources
Reproducibility is what separates a demo from a production ML system.
What to learn
- Experiment tracking (MLflow / Weights & Biases)
- Data & model versioning (DVC, model registry)
- Pipeline orchestration (Airflow / Prefect / Kubeflow)
- Feature stores (concept) & dataset lineage
Put it into practice
Your build
Track 3 training runs in MLflow and register the best model in its model registry.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
06Cloud, Containers & GPU Infrastructure4 focus areas · 2 resources
Models run on real infra. Know enough cloud + Kubernetes to deploy and scale them.
What to learn
- Docker + Kubernetes for ML workloads
- Cloud GPU instances & spot/preemptible cost control
- Managed AI platforms (Bedrock, Vertex AI, SageMaker)
- CI/CD for ML (CML, GitHub Actions for training/deploy)
Put it into practice
Your build
Deploy your inference service to a managed container platform with autoscaling.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
07Evaluation, Monitoring & Guardrails4 focus areas · 2 resources
LLMs are non-deterministic — evals and monitoring are the new unit tests.
What to learn
- Offline & online evals (LLM-as-judge, golden datasets)
- Drift detection & data/model monitoring
- Tracing LLM apps (LangSmith / OpenTelemetry GenAI)
- Safety: guardrails, PII handling, prompt-injection defense
Put it into practice
Your build
Add an eval suite + tracing to your RAG app and catch one regression with it.
Keep notes on the setup, result, and one trade-off you made.Resources & documentation
Extend the foundations
Ideas to connect as you go.
Return to these themes as the core skills become familiar.
- LLMOps: RAG, evals, prompt & context engineering
- Efficient model serving (vLLM, TGI) & GPU scheduling
- Vector databases & retrieval pipelines
- AI safety, guardrails & observability for non-determinism