1 of 66 written
Notes
Every note in study order. Work through them top to bottom.
1
Deep learning foundations
Week 1- What is a neural network? ready
- Loss functions: measuring how wrong we are to be written
- Gradient descent and backpropagation to be written
- Optimizers: SGD, Adam, AdamW and the learning rate to be written
- Logits, softmax and cross-entropy to be written
- PyTorch essentials: tensors, autograd, the training loop to be written
- Overfitting, regularization and data splits to be written
2
LLM internals: build GPT from scratch
Weeks 2–4- What is an LLM? Pretraining, fine-tuning and the big picture to be written
- Tokenization and byte-pair encoding (BPE) to be written
- Token embeddings: turning IDs into vectors to be written
- Positional encoding: telling the model about word order to be written
- Preparing data: sliding windows and next-token prediction to be written
- Self-attention, the simple version to be written
- Scaled dot-product attention: queries, keys and values to be written
- Causal masking and attention dropout to be written
- Multi-head attention to be written
- LayerNorm, GELU, feed-forward layers and residual connections to be written
- The transformer block and the full GPT architecture to be written
- Counting parameters and weight tying to be written
- Pretraining: the training loop, cross-entropy and perplexity to be written
- Decoding: greedy, temperature, top-k and top-p to be written
- Loading pretrained GPT-2 weights to be written
- Fine-tuning for classification to be written
- Instruction fine-tuning to be written
- Encoder, decoder, encoder–decoder: BERT vs GPT vs T5 to be written
3
The modern LLM stack
Week 5- From GPT-2 to modern LLMs: pre-norm, RMSNorm, SwiGLU, grouped-query attention to be written
- Rotary position embeddings (RoPE) to be written
- The KV cache: why generation is fast to be written
- Quantization: 16-bit, 8-bit, 4-bit and GGUF to be written
- LoRA, QLoRA and parameter-efficient fine-tuning to be written
- Teaching preferences: RLHF and DPO to be written
- Context windows and reasoning models to be written
- Prompting vs RAG vs fine-tuning: how to choose to be written
4
Retrieval-augmented generation (RAG)
Weeks 6–7- Text embeddings and semantic similarity to be written
- Vector databases and approximate nearest neighbour search (HNSW, IVF) to be written
- Chunking strategies to be written
- Retrieval: BM25, dense, hybrid search and rank fusion to be written
- Reranking with cross-encoders to be written
- The RAG pipeline end to end to be written
- Advanced RAG: query rewriting, HyDE, multi-hop, GraphRAG, agentic RAG to be written
- Evaluating RAG: recall@k, MRR, nDCG, faithfulness to be written
- RAG failure modes and how to debug them to be written
5
Prompting, tools and agents
Weeks 8–9- Prompt engineering that survives production to be written
- Context engineering: what goes in the window to be written
- Structured output and function (tool) calling to be written
- The agent loop and ReAct to be written
- Workflows vs agents: routing, orchestrator–workers, evaluator–optimizer to be written
- Model Context Protocol (MCP) to be written
- Agent memory: short-term, long-term, compaction to be written
- Multi-agent systems: when they help and when they hurt to be written
- Prompt injection and agent security to be written
6
Evals and production
Week 10- Evals: golden sets, metrics and regression testing to be written
- LLM-as-judge: how to trust a model grading a model to be written
- Hallucination: causes, detection, mitigation to be written
- Guardrails and safety layers to be written
- Cost and latency: caching, batching, streaming, model routing to be written
- Serving LLMs: time to first token, throughput, continuous batching to be written
- Observability and tracing to be written
7
LLM system design and interview drill
Weeks 11–12- A framework for LLM system design rounds to be written
- Design: a customer support assistant to be written
- Design: enterprise document search at scale to be written
- Design: a coding assistant to be written
- The ML coding round: attention, BPE and sampling from memory to be written
- Question bank: fundamentals to be written
- Question bank: applied AI engineering to be written
- Talking about your projects to be written