What you build

Projects

Four projects, each harder than the tutorial version, each runnable without owning a GPU.

Each project lives in its own GitHub repository, separate from these notes. Three rules apply to all four:

  1. Measure before you improve. Build the test set first, record a baseline, then change one thing at a time.
  2. Write the hard part yourself. Use frameworks only after you have built the core once by hand.
  3. The README is part of the project. It states the problem, the design, a results table and what you would do next.
1

GPT from scratch

Weeks 2–4

Write a GPT-style language model yourself, train it, then make it faster and more modern than the textbook version.

What it proves. You understand the model itself, not only how to call it.

Without a GPU. A 10–30 million parameter model trains on a laptop CPU in hours, or on a free Colab or Kaggle GPU in minutes. The real GPT-2 (124M) weights are only loaded and fine-tuned, which a free GPU handles.

Milestones

  1. Byte-pair encoding tokenizer written by hand, checked against the tiktoken library on the same text.
  2. Data loader that produces input and target pairs with a sliding window.
  3. Multi-head causal self-attention, transformer block and full GPT model in PyTorch, with a test that counts parameters.
  4. Training loop with a learning-rate warmup and cosine decay, gradient clipping and loss curves; train on the TinyStories dataset.
  5. Text generation with temperature, top-k and top-p sampling.
  6. Load OpenAI’s GPT-2 weights into your own model and confirm it generates sensible text.
  7. Fine-tune it twice: once as a spam classifier, once to follow instructions.

Stretch goals

  • Add a KV cache and measure the speed-up in tokens per second.
  • Replace learned position embeddings with rotary embeddings (RoPE) and compare the loss.
  • Swap LayerNorm and GELU for RMSNorm and SwiGLU, as modern models do.

Questions an interviewer will ask about it

  • Walk me through what happens to a sentence from text to next-token probabilities.
  • Why does your attention use a mask? What breaks without it?
  • How did you check that your implementation is correct?
2

Fine-tune, quantize and run a model locally

Week 5

Teach a small open model one real task with QLoRA, prove it improved, then shrink it so it runs on your own laptop.

What it proves. You can adapt a model cheaply and measure whether the adaptation worked.

Without a GPU. Training uses a free Kaggle or Colab T4 GPU with a model of roughly 0.5–1.5 billion parameters loaded in 4-bit. Inference runs on your CPU through llama.cpp.

Milestones

  1. Pick a narrow task with a public dataset, for example turning plain English into SQL or extracting structured fields from messy text.
  2. Build a held-out test set and an automatic scoring script before any training.
  3. Record the baseline: the untouched model, then the untouched model with a good prompt.
  4. Fine-tune with QLoRA using Hugging Face PEFT; log the loss.
  5. Score the fine-tuned model on the same test set and report the change.
  6. Merge the adapter, convert to GGUF, quantize to 4-bit and run it locally.

Stretch goals

  • Study how LoRA rank and which layers you adapt change the result.
  • Compare quality and speed at 8-bit, 5-bit and 4-bit quantization.
  • Write LoRA from scratch for your Project 1 model.

Questions an interviewer will ask about it

  • Why fine-tune here instead of prompting or RAG?
  • What does LoRA actually train, and why is that enough?
  • How do you know the model did not just memorize the training set?
3

Production RAG with evals

Weeks 6–7

A question-answering system over a large real document set, where every design choice is backed by a measured number.

What it proves. You can build the most common LLM product in industry and defend each decision in it.

Without a GPU. Hosted embedding and chat model APIs. The vector database runs locally in Docker. No GPU at any step.

Milestones

  1. Ingest a real corpus of at least a few thousand documents, for example research papers or a large open-source project’s docs.
  2. Write a golden set of 100 or more questions with the passages that answer them.
  3. Baseline: fixed-size chunks and dense retrieval only. Measure recall@k and MRR.
  4. Add BM25 and combine with reciprocal rank fusion; measure again.
  5. Add a cross-encoder reranker; measure again.
  6. Generate answers with citations; score faithfulness and answer quality with an LLM judge that you validated against your own labels.
  7. Serve it behind an API with streaming, and write up the results table.

Stretch goals

  • Query rewriting and HyDE for vague questions.
  • Contextual chunk enrichment before embedding.
  • Access control: users only retrieve documents they are allowed to see.
  • Incremental re-indexing when documents change.

Questions an interviewer will ask about it

  • How did you choose the chunk size?
  • Retrieval returns the right passage but the answer is still wrong. How do you debug that?
  • How would this change for ten million documents?
4

An agent with your own MCP server, run like a production service

Weeks 8–10

An agent that completes multi-step tasks with tools you wrote, plus everything needed to trust it: evals, tracing, guardrails and a cost report.

What it proves. You can ship an agent and show, with data, how often it succeeds, what it costs and how it fails.

Without a GPU. Hosted model APIs only.

Milestones

  1. Pick a real job, for example an agent that investigates a failing test in a repository, or one that answers questions over a database.
  2. Write the agent loop yourself first, with no framework: model call, tool call, result, repeat.
  3. Build an MCP server that exposes your tools, and connect the agent to it.
  4. Create 30 or more test tasks with checkable outcomes; report the success rate.
  5. Add tracing so every step, token count and tool call of a run can be inspected.
  6. Add guardrails: input checks, tool permission limits, a step budget and defence against prompt injection in tool results.
  7. Cut cost and latency with prompt caching and a cheaper model for easy steps; report before and after.

Stretch goals

  • Long-running tasks with context compaction and memory.
  • Compare a single agent against an orchestrator with sub-agents on the same test tasks.
  • Human approval step for risky actions.

Questions an interviewer will ask about it

  • When would you use a fixed workflow instead of an agent?
  • Your agent succeeds 70% of the time. How do you find out why it fails the rest?
  • How do you stop a malicious document from hijacking the agent?

Putting them on a resume

A project line has three parts: what you built, the hard technical choice, and a number you measured yourself.