Phase 1 · Week 1 · Note 1 of 7
What is a neural network?
A neural network is a big adjustable formula. This note builds it up from its smallest part, the neuron, to a full network, with worked numbers at every step.
Why this note matters
Everything later in these notes is built from the idea on this page. A large language model such as GPT is a neural network. Attention, embeddings, fine-tuning and LoRA are all ways of arranging or adjusting the same basic part: a layer of numbers that multiplies its input and bends the result.
Interviewers rarely ask “what is a neural network?” to an experienced candidate, but they constantly ask questions that depend on it: why activation functions are needed, how many parameters a layer has, what shape a tensor is after a layer, why weights cannot start at zero. If this page is solid, those are easy marks.
Before reading on, watch this. It gives you the pictures; this note gives you the words and the detail.
Watch 3Blue1Brown: But what is a neural network? (Deep learning, chapter 1) About 19 minutes. The best visual introduction there is. Watch it once fully before the rest of this note.In plain words
Each everyday idea has an exact technical name. You need both: the first to understand, the second to speak.
| Everyday idea | Technical term | Symbol |
|---|---|---|
| The things you look at | inputs or features | |
| How much each one matters | weights | |
| Your general willingness to say yes | bias | |
| The total of all the pushes | pre-activation (weighted sum) | |
| “Is the total high enough?” | activation function | |
| Your decision | output, or activation | |
| All the “how much it matters” numbers | parameters | |
| Finding good numbers from examples | training |
A worked example
Let us put real numbers on the walk decision. Each input is 1 for yes and 0 for no. Suppose these are the weights and the bias:
| Input | Weight | Meaning |
|---|---|---|
| Sunny | +3 | Sunshine pushes strongly towards going. |
| Cold | −2 | Cold pushes against going. |
| Tired | −4 | Tiredness pushes against going even more. |
| Bias | −1 | With no information at all, you lean slightly towards staying home. |
The rule: add everything up, and go if the total is above zero.
Day A: sunny, not cold, not tired.
Day B: sunny, cold and tired.
That is a complete neuron. Notice three things. A positive weight makes an input argue for “yes”, a negative weight makes it argue for “no”, and the size of the weight is how loudly it argues. The bias moves the bar: with a bias of +5 you would go on almost any day; with −5 almost never.
Here we chose the weights by hand. The whole point of machine learning is that we do not. We show the network many days along with what the right decision was, and an algorithm adjusts the weights until the outputs match. How that adjustment works is the subject of the next two notes.
One neuron, precisely
A neuron does two steps.
- Weighted sum. Multiply every input by its weight, add the results, add the bias. The total is the pre-activation, written . Mathematically this is a dot product of the weight vector and the input vector, plus the bias.
- Activation. Pass through a fixed function . The result is the neuron’s output, and becomes an input to neurons in the next layer.
In the walk example the activation was a hard rule: output 1 if , otherwise 0. That is called a step function, and a neuron that uses it is the original perceptron, invented by Frank Rosenblatt in 1958. Modern networks replace the hard step with smoother functions, for a reason that matters a lot: training works by nudging weights and watching whether the output gets slightly better. A step function is flat everywhere, so a small nudge changes nothing and gives no signal about which direction to move. Smooth activations have a slope, and the slope is the signal.
What one neuron can and cannot do
Look again at the rule “go if ”. The points where the sum is exactly zero form a straight line when there are two inputs, a flat plane with three, and a hyperplane in general. Everything on one side gets “yes”, everything on the other side gets “no”. This dividing surface is the decision boundary.
So a single neuron can only separate things that one straight cut can separate. Data like that is called linearly separable. The weights set the tilt of the line and the bias slides it sideways.
The classic example of what one neuron cannot do is XOR (“exclusive or”): output 1 when exactly one of two inputs is 1. In Fig. 2 the two “yes” dots sit on opposite corners, and no single straight line puts them on one side and the two “no” dots on the other.
This is not a trivia point. In 1969 Marvin Minsky and Seymour Papert showed this limitation in their book Perceptrons, and interest in neural networks collapsed for years. The way out, which you will see worked through below, is to stack neurons in layers with a non-linear activation between them.
Read Michael Nielsen: Neural Networks and Deep Learning, chapter 1 Read the sections “Perceptrons” and “Sigmoid neurons”. They cover this exact step, from hard step to smooth activation, slowly and clearly.Why the bend matters
The weighted sum is a linear operation. If you stack linear operations with nothing in between, the result is still one linear operation. Take two layers with no activation and substitute the first into the second:
The right-hand side has the form “one matrix times , plus one vector”. That is a single layer with weights and bias . The same argument repeats for any number of layers, so a hundred-layer network without activations is exactly as powerful as one layer, and still cannot solve XOR.
The activation function adds non-linearity. Once each layer bends its output, the next layer is combining bent pieces, and stacking them builds shapes of any complexity. This is the single most common “why” question about neural networks, so learn the collapse argument well enough to write it on a whiteboard.
The activation functions to know
You should be able to name each of these, sketch its shape, and state one strength and one weakness. The “slope” column matters because training multiplies slopes together across layers; you will see why in the backpropagation note.
| Name | Formula | Output range | Slope | Where it is used | Weakness |
|---|---|---|---|---|---|
| Step | 1 if , else 0 | 0 or 1 | 0 everywhere | The 1958 perceptron | No slope, so it cannot be trained with gradients. |
| Sigmoid | 0 to 1 | At most 0.25 | Output layer for yes/no probability; gates inside LSTMs | Flat at both ends (it saturates), which causes vanishing gradients. | |
| Tanh | −1 to 1 | At most 1 | Older recurrent networks | Still saturates at both ends. | |
| ReLU | 0 to ∞ | 0 or exactly 1 | Default for most networks | “Dying ReLU”: a neuron stuck below zero outputs 0 and stops learning. | |
| GELU | about −0.17 to ∞ | Smooth, near 1 for large z | GPT-2, BERT, most transformers | Slightly more expensive to compute. | |
| SiLU / Swish | about −0.28 to ∞ | Smooth, near 1 for large z | Inside SwiGLU in Llama-style models | Same as GELU. |
In the GELU row, is the standard normal cumulative distribution: the probability that a standard bell-curve value is below . You can read GELU as “multiply by how confident we are that it is large”. For big positive it behaves like ReLU; near zero it curves gently instead of making a sharp corner.
Paper Hendrycks and Gimpel: Gaussian Error Linear Units (GELUs), 2016 Short paper. Read section 2 for the definition and Figure 1 for the comparison with ReLU.One more activation deserves a mention here, although it gets its own note: softmax. It is used on the output layer when the network must choose among many options, and it turns a list of scores into probabilities that add up to 1. An LLM’s final step is a softmax over every token in its vocabulary.
From neuron to network
Put several neurons side by side, all reading the same inputs, and you have a layer. Feed one layer’s outputs into the next and you have a network. When every neuron connects to every output of the layer before it, the layer is called fully connected (also “dense” or “linear”), and the whole thing is a multilayer perceptron (MLP) or feed-forward network. “Feed-forward” means data flows one way, from input to output, with no loops.
- The input layer is just the data, written as a list of numbers (a vector). It does no computing and has no parameters.
- Hidden layers sit in the middle. They are “hidden” because their values are neither the input you provide nor the output you read; the network decides for itself what they should represent.
- The output layer gives the answer, for example one score per possible class. The number of neurons in a layer is its width; the number of layers is the network’s depth.
Running data through the network from input to output is the forward pass. It is what happens every time you use a trained model (called inference), and it is also the first half of every training step.
What do hidden neurons end up representing? Nobody assigns them a job. During training each one drifts towards detecting some pattern that helps reduce the error. In an image network, first-layer neurons typically respond to edges, later ones to textures and shapes, and the last ones to whole objects. Learning these intermediate descriptions automatically is called representation learning, and it is the main reason deep learning replaced hand-written features.
Try it TensorFlow Playground Pick the circle dataset. Train with zero hidden layers and watch it fail, then add one hidden layer of 3 neurons. Hover over each neuron to see what it learned. Ten minutes here is worth an hour of reading.A hidden layer solves XOR
Here is the promised proof that layers plus a bend beat a single neuron. This tiny network has two inputs, two hidden neurons with ReLU, and one output:
Run all four inputs through it by hand:
| XOR | |||||
|---|---|---|---|---|---|
| 0 | 0 | ReLU(0) = 0 | ReLU(−1) = 0 | 0 | 0 ✓ |
| 0 | 1 | ReLU(1) = 1 | ReLU(0) = 0 | 1 | 1 ✓ |
| 1 | 0 | ReLU(1) = 1 | ReLU(0) = 0 | 1 | 1 ✓ |
| 1 | 1 | ReLU(2) = 2 | ReLU(1) = 1 | 2 − 2 = 0 | 0 ✓ |
How it works: counts how many inputs are on. stays at zero until both are on, and that is only possible because ReLU cuts off the negative value in the first row. The output then subtracts twice to cancel the “both on” case. Remove the ReLU and would be −1, 0, 0, 1, the trick disappears, and you are back to a single straight line.
The hidden layer has changed the problem. In the original inputs the four points could not be separated by a line. In the new coordinates they can. That is the general picture of what every hidden layer does: it re-describes the data so that the next layer’s job is easier.
Matrices and shapes
Writing each neuron separately does not scale, so a whole layer is written as one matrix operation. Stack each neuron’s weights as one row of a weight matrix , and collect the biases into a vector :
For the hidden layer of Fig. 4, with 3 inputs and 4 neurons, the pieces have these shapes:
Row of holds the weights of neuron , so the matrix has one row
per output and one column per input: shape (out_features, in_features). The activation is then
applied to each of the four numbers separately (element-wise).
In practice we never process one example at a time. We stack many examples into a
batch, giving the input a shape of (batch_size, in_features), and the layer
processes all rows at once. The output has shape (batch_size, out_features). A linear layer only ever
changes the last dimension. This is true in a transformer too, where the input is
(batch, sequence_length, embedding_size) and a linear layer transforms just the final number.
This is also why deep learning needs GPUs. The work is almost entirely multiplying big matrices, which is thousands of independent multiply-and-add operations, and a GPU is hardware built to do those in parallel.
Read PyTorch documentation: torch.nn.Linear Two minutes. Check the “Shape” section and the shapes of .weight and .bias against what you just read.Counting parameters
A fully connected layer with inputs and neurons has weights and biases. Activation functions such as ReLU and GELU have no parameters. For the network in Fig. 4:
When people say GPT-2 small has 124 million parameters, they mean the same thing: 124 million individual numbers like these, all adjusted during training. Being able to do this count quickly is a real interview skill. It shows up as “how big is this layer?”, “how much memory does this model need?” and, later, “how many parameters does LoRA train?”
A useful conversion: each parameter stored as a 32-bit float takes 4 bytes, so a model’s size in memory is roughly 4 bytes × the parameter count. For 124 million parameters that is about 0.5 GB; at 16-bit precision, half that.
Do not confuse parameters with hyperparameters. Parameters are learned from data. Hyperparameters are choices you make before training: how many layers, how wide, which activation, what learning rate.
Why deep, not just wide
A famous result, the universal approximation theorem, says that a network with just one hidden layer and a suitable non-linear activation can approximate any continuous function on a bounded region as closely as you like, provided the hidden layer is wide enough.
That sounds as if depth is unnecessary. It is not, because of what the theorem leaves out:
- It does not say how many neurons “enough” is. For many functions the number is impractically huge.
- It says such weights exist, not that training will find them.
- It says nothing about performing well on new data the network has not seen.
Depth solves the first problem. A deep network builds features in a hierarchy, reusing simple features to make complex ones, the way words are reused to build many sentences. For some functions a shallow network needs exponentially more neurons than a deep one to do the same job. In practice, deep networks also train to better solutions on structured data such as images and language.
Read Michael Nielsen: A visual proof that neural nets can compute any function Interactive. You drag weights and watch one hidden layer build any curve out of small steps. Makes the theorem obvious without any math.How the weights start
Before training, the weights have to be set to something. This is initialization, and there is one rule you must know: never start all weights at the same value, including zero.
Suppose every weight in a layer is zero. Every neuron in that layer then computes exactly the same output for any input. When training asks “how should each neuron change?”, the answer is identical for all of them, so they all change by the same amount and remain identical forever. You have paid for a hundred neurons and got one. Starting with small random values makes each neuron slightly different so they can learn different things. This is called symmetry breaking.
The size of the random values matters too. Too large, and signals grow layer after layer until they overflow. Too small, and they shrink to nothing. Two standard recipes pick the scale from the layer’s size so that signals keep roughly the same spread from layer to layer:
- Xavier (Glorot) initialization, designed for sigmoid and tanh. Paper: Glorot and Bengio, 2010.
- He (Kaiming) initialization, designed for ReLU. Paper: He et al., 2015.
Biases are normally started at zero, which is fine because the random weights already break the symmetry. GPT-2 initializes its weights from a normal distribution with a standard deviation of 0.02.
Code
Type these out yourself rather than pasting. Before running each one, write down what you expect it to print.
One neuron in NumPy
This is Fig. 1 with numbers: a dot product, a bias, and ReLU.
import numpy as np
x = np.array([1.0, 2.0, 3.0]) # three inputs
w = np.array([0.5, -1.0, 2.0]) # one weight per input
b = 0.1 # bias
z = np.dot(w, x) + b # weighted sum: 0.5 - 2.0 + 6.0 + 0.1 = 4.6
a = max(0.0, z) # ReLU activation: 4.6
print(z, a)
The XOR network in NumPy
The same network as the hand-worked table above, now in matrix form. X @ W1.T is the matrix
multiplication that computes every neuron’s weighted sum for all four inputs at once. The transpose
.T is needed because each row of W1 is one neuron’s weights and each row of
X is one example.
import numpy as np
X = np.array([[0, 0], [0, 1], [1, 0], [1, 1]], dtype=float) # all four inputs, one per row
W1 = np.array([[1, 1], # weights of hidden neuron h1
[1, 1]]) # weights of hidden neuron h2
b1 = np.array([0, -1]) # biases of h1 and h2
W2 = np.array([[1, -2]]) # output neuron: y = 1*h1 - 2*h2
b2 = np.array([0])
H = np.maximum(0, X @ W1.T + b1) # hidden layer with ReLU, shape (4, 2)
Y = H @ W2.T + b2 # output layer, shape (4, 1)
print(Y.ravel()) # [0. 1. 1. 0.] -> XOR
The network of Fig. 4 in PyTorch
import torch
import torch.nn as nn
model = nn.Sequential(
nn.Linear(3, 4), # 3 inputs -> 4 hidden neurons: 3*4 weights + 4 biases = 16
nn.ReLU(), # the bend; has no parameters
nn.Linear(4, 2), # 4 hidden -> 2 outputs: 4*2 weights + 2 biases = 10
)
x = torch.tensor([[1.0, 2.0, 3.0]]) # shape (1, 3): a batch of one example
y = model(x) # the forward pass
print(y.shape) # torch.Size([1, 2])
print(model[0].weight.shape) # torch.Size([4, 3]) -> (out_features, in_features)
print(model[0].bias.shape) # torch.Size([4])
print(sum(p.numel() for p in model.parameters())) # 26
nn.Linear(3, 4) stores a weight matrix of shape (4, 3) and a bias of shape
(4,), and computes . The weights start as small random
numbers, so the output is meaningless until the network is trained.
Where this lives in an LLM
None of this is just history. Every transformer block inside GPT contains a two-layer feed-forward network of exactly the kind on this page: a linear layer, a GELU activation, and another linear layer. In GPT-2 small the first of those layers goes from 768 numbers to 3,072, which is parameters in one layer.
Within each transformer block, these feed-forward layers hold roughly two thirds of the weights, and attention holds the other third. So when you study the GPT architecture in Phase 2, a large part of each block will already be familiar. The Vizuara course builds this block in code, and you can follow along from the start of the playlist:
Watch Vizuara: Building LLMs from scratch (full playlist) The main course for Phase 2. The first few lectures are an overview and can be watched this week alongside these foundation notes.The full interview answer
The short blue lines above are what to say for each sub-topic. This is how they join into one answer to “explain what a neural network is”. Aim for about ninety seconds.
If they push further, be ready for:
- “Show me why it collapses.” Write the two-layer substitution from this section.
- “How many parameters?” Inputs × outputs + outputs, per layer.
- “Why not initialize to zero?” Symmetry: identical neurons get identical updates.
- “Why ReLU over sigmoid?” Sigmoid’s slope is at most 0.25 and it saturates, so gradients vanish in deep networks.
Practice questions
Answer each one out loud before opening it.
Why do neural networks need activation functions?
To introduce non-linearity. A composition of linear functions is itself linear, so without activations a deep network is equivalent to a single linear layer and can only model straight-line relationships. Non-linear activations let stacked layers represent curved, complex decision boundaries.
Can a single neuron learn XOR? Why or why not?
No. A single neuron’s decision boundary is a hyperplane, so it can only separate linearly separable classes. In XOR the two positive points lie on opposite corners of the square and no straight line separates them from the two negative points. Adding one hidden layer with a non-linear activation solves it, because the hidden layer maps the inputs into a space where they are linearly separable.
What is the difference between a weight and a bias?
A weight scales one input and controls how strongly that input influences the neuron. The bias is a constant added to the weighted sum that does not depend on any input; it shifts the activation threshold. Geometrically, weights set the orientation of the decision boundary and the bias moves it away from the origin. Without a bias the boundary would be forced to pass through the origin.
What is the difference between a parameter and a hyperparameter?
Parameters are learned from data during training: the weights and biases. Hyperparameters are chosen by the engineer before training and are not learned: the number of layers, the layer width, the learning rate, the batch size, the choice of activation.
How many parameters does a linear layer from 768 to 3,072 units have?
Weights: 768 × 3,072 = 2,359,296. Biases: 3,072. Total: 2,362,368. The general rule is inputs × outputs + outputs. This particular layer is the first half of the feed-forward block in GPT-2 small.
An input of shape (32, 10, 768) goes through nn.Linear(768, 3072). What is the output shape, and what is the shape of the weight?
The output is (32, 10, 3072): a linear layer transforms only the last dimension and leaves the others alone. The weight has shape (3072, 768), which is (out_features, in_features), and the bias has shape (3072,).
Why is ReLU usually preferred to sigmoid in hidden layers?
Sigmoid saturates: for large positive or negative inputs its slope is almost zero, and its slope is never more than 0.25. During backpropagation those small slopes multiply across layers and the gradient shrinks towards zero, the vanishing gradient problem, so early layers stop learning. ReLU has a slope of exactly 1 for positive inputs, which keeps gradients alive, and it is cheaper to compute. Its weakness is “dying ReLU”: a neuron whose input is always negative outputs zero and gets no gradient. Variants such as Leaky ReLU and GELU address this.
Why can’t we use a step function as the activation in a modern network?
Its derivative is zero everywhere (and undefined at the jump), so gradient-based training gets no signal about how to change the weights. Training needs activations that are differentiable, or at least have a useful slope almost everywhere, which is why smooth or piecewise-linear functions replaced it.
What happens if you initialize all weights to zero?
Every neuron in a layer computes the same output and receives the same gradient, so they all update identically and stay identical. The layer behaves like a single neuron regardless of its width. Random initialization breaks this symmetry. Biases can safely start at zero.
What does the universal approximation theorem say, and what does it not say?
It says a feed-forward network with a single hidden layer and a suitable non-linear activation can approximate any continuous function on a bounded input region as closely as you like, given enough hidden neurons. It does not say how many neurons are needed (possibly an impractical number), nor that training will actually find the right weights, nor that the result will generalize to new data.
If one wide layer is enough in theory, why do we use deep networks?
Depth is far more parameter-efficient. Deep networks build features in a hierarchy, reusing simple features to form complex ones, so for some functions they need exponentially fewer neurons than a shallow network. In practice deep networks also generalize better on structured data such as images and language.
What is a forward pass?
Computing the network’s output for a given input by applying each layer in order, from the input layer to the output layer. It is what happens at inference time, and it is the first half of every training step; the second half is the backward pass, which computes gradients.
Roughly how much memory do the weights of a 124-million-parameter model need?
At 32-bit floating point each parameter takes 4 bytes, so about 124M × 4 ≈ 496 MB, roughly 0.5 GB. At 16-bit precision it is about 0.25 GB. This is weights only; running the model also needs memory for activations.
Common mistakes
All sources
Everything linked above, in the order to use it, plus two deeper references.
- 3Blue1Brown, “But what is a neural network?” (video). Start here.
- StatQuest, “The Essential Main Ideas of Neural Networks” (video). A second, slower explanation if the first did not click.
- Michael Nielsen, Neural Networks and Deep Learning, chapter 1 and chapter 4 (free online book).
- TensorFlow Playground (interactive).
- Andrej Karpathy, micrograd (video, code).
- Sebastian Raschka, Build a Large Language Model (From Scratch), Appendix A, with its code; chapter 4 for the feed-forward block and GELU.
- Goodfellow, Bengio and Courville, Deep Learning, chapter 6: “Deep Feedforward Networks” (free online). The rigorous version, including the XOR example.
- Universal approximation theorem (Wikipedia), for the precise statements and their history.
- Glorot, Bordes and Bengio, “Deep Sparse Rectifier Neural Networks”, 2011. An influential paper showing ReLU works well in deep networks.