The learning path
Roadmap
Twelve weeks, seven phases, four projects. Internals first, then building real systems on top.
How to study each day
The order matters. Meeting an idea three times, in three forms, is what makes it stick.
- Watch or read the source for the topic (course lecture or book chapter listed under the phase).
- Read our note on the same topic. Redraw its main diagram on paper from memory.
- Type the code yourself. Do not paste. Change one thing and predict what happens before you run it.
- Say the interview answer out loud without looking. If you get stuck, that is the part to reread.
- Answer the practice questions before opening them.
| Day | What to do |
|---|---|
| Monday to Friday | One or two notes a day, following the five steps above. |
| Saturday | Project work only. Build the milestone for that week. |
| Sunday | Revision cards and practice questions for the whole week, then rest. |
The seven phases
Deep learning foundations
Week 1Goal. Understand how any neural network learns, and be able to write a training loop in PyTorch from memory.
Study from.
- 3Blue1Brown, “Neural Networks” series (chapters 1–4)
- Andrej Karpathy, “Zero to Hero”: the micrograd video
- Raschka, Build a Large Language Model (From Scratch), Appendix A: Introduction to PyTorch
Notes.
- What is a neural network? ready
- Loss functions: measuring how wrong we are to be written
- Gradient descent and backpropagation to be written
- Optimizers: SGD, Adam, AdamW and the learning rate to be written
- Logits, softmax and cross-entropy to be written
- PyTorch essentials: tensors, autograd, the training loop to be written
- Overfitting, regularization and data splits to be written
LLM internals: build GPT from scratch
Weeks 2–4Goal. Explain and code every part of a GPT-style model: text in, next token out.
Build. Project 1: GPT from scratch
Study from.
- Vizuara, “Building LLMs from Scratch” (the full playlist, at 1.5–2x)
- Raschka, Build a Large Language Model (From Scratch), chapters 1–7
- Vaswani et al., “Attention Is All You Need” (2017)
- Jay Alammar, “The Illustrated Transformer” and “The Illustrated GPT-2”
- Karpathy, “Let’s build GPT: from scratch, in code, spelled out”
Notes.
- What is an LLM? Pretraining, fine-tuning and the big picture to be written
- Tokenization and byte-pair encoding (BPE) to be written
- Token embeddings: turning IDs into vectors to be written
- Positional encoding: telling the model about word order to be written
- Preparing data: sliding windows and next-token prediction to be written
- Self-attention, the simple version to be written
- Scaled dot-product attention: queries, keys and values to be written
- Causal masking and attention dropout to be written
- Multi-head attention to be written
- LayerNorm, GELU, feed-forward layers and residual connections to be written
- The transformer block and the full GPT architecture to be written
- Counting parameters and weight tying to be written
- Pretraining: the training loop, cross-entropy and perplexity to be written
- Decoding: greedy, temperature, top-k and top-p to be written
- Loading pretrained GPT-2 weights to be written
- Fine-tuning for classification to be written
- Instruction fine-tuning to be written
- Encoder, decoder, encoder–decoder: BERT vs GPT vs T5 to be written
The modern LLM stack
Week 5Goal. Know what changed between GPT-2 and today’s models, and how models are adapted and shrunk cheaply.
Build. Project 2: QLoRA fine-tune, quantize, run locally
Study from.
- Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models” (2021)
- Dettmers et al., “QLoRA” (2023)
- Su et al., “RoFormer: Rotary Position Embedding” (2021)
- Raschka, Build a Large Language Model (From Scratch), Appendix E: LoRA
- Hugging Face PEFT and Transformers documentation
Notes.
- From GPT-2 to modern LLMs: pre-norm, RMSNorm, SwiGLU, grouped-query attention to be written
- Rotary position embeddings (RoPE) to be written
- The KV cache: why generation is fast to be written
- Quantization: 16-bit, 8-bit, 4-bit and GGUF to be written
- LoRA, QLoRA and parameter-efficient fine-tuning to be written
- Teaching preferences: RLHF and DPO to be written
- Context windows and reasoning models to be written
- Prompting vs RAG vs fine-tuning: how to choose to be written
Retrieval-augmented generation (RAG)
Weeks 6–7Goal. Design, build and measure a RAG system, and explain every trade-off in it.
Build. Project 3: production RAG with evals
Study from.
- Chip Huyen, AI Engineering (2025), the RAG and agents chapter
- Jay Alammar and Maarten Grootendorst, Hands-On Large Language Models
- Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks” (2020)
- Malkov and Yashunin, “HNSW” (2016)
- Anthropic, “Introducing Contextual Retrieval” (2024)
Notes.
- Text embeddings and semantic similarity to be written
- Vector databases and approximate nearest neighbour search (HNSW, IVF) to be written
- Chunking strategies to be written
- Retrieval: BM25, dense, hybrid search and rank fusion to be written
- Reranking with cross-encoders to be written
- The RAG pipeline end to end to be written
- Advanced RAG: query rewriting, HyDE, multi-hop, GraphRAG, agentic RAG to be written
- Evaluating RAG: recall@k, MRR, nDCG, faithfulness to be written
- RAG failure modes and how to debug them to be written
Prompting, tools and agents
Weeks 8–9Goal. Build reliable LLM features: good context, structured output, tool use and agent loops.
Build. Project 4: agent with your own MCP server
Study from.
- Anthropic, “Building Effective Agents” (2024)
- Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models” (2022)
- Model Context Protocol specification and documentation
- Anthropic and OpenAI prompt engineering guides
- Chip Huyen, AI Engineering (2025), the prompt engineering chapter
Notes.
- Prompt engineering that survives production to be written
- Context engineering: what goes in the window to be written
- Structured output and function (tool) calling to be written
- The agent loop and ReAct to be written
- Workflows vs agents: routing, orchestrator–workers, evaluator–optimizer to be written
- Model Context Protocol (MCP) to be written
- Agent memory: short-term, long-term, compaction to be written
- Multi-agent systems: when they help and when they hurt to be written
- Prompt injection and agent security to be written
Evals and production
Week 10Goal. Prove an LLM feature works, then run it cheaply, quickly and safely.
Build. Project 4 continued: evals, guardrails, cost and latency report
Study from.
- Chip Huyen, AI Engineering (2025), the evaluation and inference optimization chapters
- Hamel Husain, “Your AI Product Needs Evals”
- vLLM documentation (continuous batching, PagedAttention)
Notes.
- Evals: golden sets, metrics and regression testing to be written
- LLM-as-judge: how to trust a model grading a model to be written
- Hallucination: causes, detection, mitigation to be written
- Guardrails and safety layers to be written
- Cost and latency: caching, batching, streaming, model routing to be written
- Serving LLMs: time to first token, throughput, continuous batching to be written
- Observability and tracing to be written
LLM system design and interview drill
Weeks 11–12Goal. Turn everything into spoken, structured answers under time pressure.
Build. Polish all four projects: READMEs, demos, resume bullets
Study from.
- Our own notes, phases 1–6
- Alex Xu, System Design Interview (for the classic distributed-systems parts)
- Chip Huyen, AI Engineering (2025), the architecture chapter
Notes.
- A framework for LLM system design rounds to be written
- Design: a customer support assistant to be written
- Design: enterprise document search at scale to be written
- Design: a coding assistant to be written
- The ML coding round: attention, BPE and sampling from memory to be written
- Question bank: fundamentals to be written
- Question bank: applied AI engineering to be written
- Talking about your projects to be written
How you know a phase is done
- You can redraw every main diagram of the phase on a blank page.
- You can give each note’s interview answer in under two minutes, without notes.
- You can answer at least four in five practice questions correctly.
- The project milestone for the phase runs, and its README explains what you measured.