The lesson in one minute
What you'll be able to explain
- A neuron is a weighted sum plus bias through a nonlinearity; a layer is a matrix multiply; a network is stacked layers.
- Without nonlinearity, depth is pointless: stacked linear layers equal one.
- Training loop: forward, loss, backprop (the chain rule gives a gradient per weight), optimizer step. Repeat.
- Backprop reuses forward activations, so training needs far more memory than inference.
Level 1
The practitioner's guide
In one sentence
A neural network is stacked layers of weighted sums with a nonlinear rule after each, and training is one loop, forward pass, loss, backpropagation, optimizer step, that nudges every weight in the direction that makes the loss smaller.
When you need it
You need this the first time you run or read a
training job rather than call a finished model: a fine-tuning API asks you
for epochs and a batch size, a training script runs out of memory that
inference never needed, a loss curve refuses to move, or a custom layer
needs a gradient you must trust. The tell: a setting in a training config
(num_train_epochs, per_device_train_batch_size, hidden_act) that you
copied from an example without knowing which box of the loop it touches.
You don't need it to prompt a hosted model, and you don't need it to pick
one from a leaderboard; the forward pass is the only part that runs at
inference and a vendor has already tuned it. One number from this lesson
says why depth is worth anything: on the two-moons data, a single linear
layer scores 88% because it can only draw a straight line, and one hidden
layer with a nonlinearity scores 99.75% (one point of 400 wrong).
Your options
How much of the training loop you own, from the least to the most:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Call a trained model | Runs the forward pass only | No gradients, no activations kept: inference memory, and nothing you can break | Per-token pricing, and a model whose weights you cannot move | The API or the model server |
| Hosted fine-tuning | The vendor runs forward, loss, backward and step on your examples; you set epochs and batch size | No training infrastructure; the loop's four boxes are the vendor's problem | Curated examples, a job that takes hours, a higher per-token price at some vendors | The vendor's fine-tuning API |
| Framework autograd on your own hardware | loss.backward() in PyTorch or jax.grad records the forward pass and replays the chain rule backwards over every weight |
The gradient is right, for any layer you compose from the framework's pieces | GPU memory several times the inference footprint, because every layer's activations are cached for the backward pass; the loop is yours to debug | Your training script |
| Backprop by hand, checked numerically | You write the backward pass of a custom layer and compare it with nudging each weight by ±ε | A gradient you can prove: this lesson's check agrees to a relative error of 4.5 × 10⁻⁹ | Time, and a test per layer | A custom layer's backward method |
How to choose
Start from what you are changing.
- Behaviour that prompting cannot fix, on a model you cannot host: hosted fine-tuning. Keep the epochs low; fine-tuning runs several passes over a small dataset, and that is how it overfits.
- A model you host and a dataset of your own: the framework, with the
activation the model was built with. Read it from the config
(
hidden_actin a Hugging Face Llama config issilu; GPT-2 uses GELU; the transformer paper's feed-forward uses ReLU) and leave it alone, because the trained weights assume it. - A new layer, a new loss, a new attention variant: write its backward pass and gradient-check it before training anything on it. A wrong gradient trains quietly and badly.
- A training run that misbehaves: the fault is in one of the four boxes (data, forward, loss, backward-and-step). Plot the loss per epoch first; its shape (slow start, steep fall, flat tail) is the first diagnostic in this lesson.
- Whatever you pick, hold out an evaluation set. The loss you train on is the loss the loop minimises, not the score you care about.
What it costs
Memory first: the backward pass reuses every layer's
activations from the forward pass, so training keeps what inference throws
away and needs several times the memory (this lesson's two-layer trace
shows the three cached values it depends on). Compute is one matrix
multiply per layer per direction, and in a transformer most parameters sit
in the feed-forward layers, which are exactly the two-layer network built
here. Steps are counted in batches: 400 examples in batches of 32 is 13
steps per epoch, and two epochs is 26 updates; Hugging Face's
TrainingArguments defaults to 3 epochs, a batch of 8 per device and a
learning rate of 5 × 10⁻⁵. Pretraining a language model is roughly one
epoch over trillions of tokens; fine-tuning is several epochs over a small
set. Quality is paid for in gradient signal: a sigmoid's slope is at most
0.25, so ten sigmoid layers can shrink the gradient by 0.25¹⁰ ≈ 10⁻⁶ before
it reaches the first weights, while ReLU passes exactly 1 for positive
inputs.
What breaks
- The loss doesn't move. The learning rate is too small (the one-knob example halves the error every step at 0.25 and would crawl at 0.001), or the gradient is dying on the way back: saturated sigmoids pass 0.02 of the signal at |z| = 4, and a ReLU whose input stays negative passes nothing, forever.
- The loss explodes or turns NaN. The step is too big;
primer.ml.optimizersis the fix. - Out of memory in training, fine in inference. Cached activations. Halve the batch size before anything else.
- Great training loss, poor real results. Several epochs over a small dataset memorised it; measure on held-out data and stop earlier.
- A custom gradient that is wrong. The network still trains, just not towards anything. Nudge each weight by ±ε and compare: anything below about 10⁻⁶ relative error is right, this lesson's check reports 4.5 × 10⁻⁹.
- Depth without nonlinearity. Five stacked linear layers equal one matrix to 2.2 × 10⁻¹⁶; more parameters, no more power. Every layer needs its activation.
In the wild
PyTorch's autograd does this lesson's backward pass for
billions of weights when you call loss.backward(); JAX does it as a
function transformation, jax.grad and jax.value_and_grad. PyTorch's
nn.Linear initialises its weights from a uniform range set by the
fan-in, the rule primer.ml.deep_nets explains. Activations in the field:
GELU in BERT and GPT-2 (Hendrycks and Gimpel, 2016), SwiGLU in Llama 2
(its paper), and hidden_act is the field a Hugging Face config uses to
name it. Hugging Face's Trainer runs the four-box loop with the defaults
above, and any hosted fine-tuning API runs the same loop behind a form that
asks for epochs and a batch size. The idea is Rumelhart, Hinton and
Williams (1986); Karpathy's micrograd is the whole of autograd in about a
hundred lines. The papers are linked at the end of the lesson.
Go deeper
Level 2 trains one knob by hand, builds a neuron and a layer, passes blame backwards through the chain rule with every number shown, proves why stacked linear layers collapse, compares the activations and their slopes, and runs backprop through two layers against a numerical check. If you only needed to read a training config, you are done.
Level 2
How it works, from scratch
A neuron is six lines of arithmetic, a layer is a matrix multiply, and training is a loop of four boxes. This level builds each one with numbers small enough to check by hand, then verifies the hand-written gradients against a numerical check.
Chapter 1
The idea: learning is adjusting knobs to be less wrong
Picture learning to throw darts blindfolded, with a friend calling out "a bit high, a bit left". You throw, hear how far off you were, and adjust your arm a little. After enough throws your arm "knows" where the bullseye is. A neural network learns the same way. Its "arm" is millions of numbers called weights, and the friend is a formula (the loss) that says how far off each guess was and which way to adjust.
Two words carry this whole lesson. The derivative (or slope) of the
loss with respect to a weight answers "if I nudge this weight up a tiny bit,
does the loss go up or down, and how fast?" The gradient is simply the
list of those slopes, one per weight. Stepping against the gradient is
walking downhill. (New to the notation? primer.notation builds every symbol
used here from zero.)
Worked example with a single knob: a model with one weight w predicts y = w·x and should learn y = 3x from the example x = 1, y = 3. The loss is (w·1 − 3)². Its slope with respect to w is 2(w − 3). Start at w = 0 and step against the slope with step size 0.25:
| step | w | slope 2(w − 3) | new w = w − 0.25·slope |
|---|---|---|---|
| 0 | 0 | −6 | 1.5 |
| 1 | 1.5 | −3 | 2.25 |
| 2 | 2.25 | −1.5 | 2.625 |
Each step halves the distance to 3. That's all training is.
Figure 1 · Diagram
flowchart LR D[Batch of examples] --> F[Forward pass<br/>make predictions] F --> L[Loss<br/>how wrong were we?] L --> B[Backpropagation<br/>gradient per weight] B --> O[Optimizer step<br/>adjust weights] O -->|next batch| D
The update rule, for every weight at once, with learning rate η:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| one weight (a knob the model can turn) | 0 at the start | |
| "replace the old value with" (an update, not an equation) | ||
| "eta", the learning rate: how big a step to take | 0.25 | |
| the loss: one number saying how wrong the model is | ||
| the slope of the loss with respect to (the "partial derivative": how the loss changes when only moves) |
In words: "the new weight is the old weight minus the learning rate times the slope of the loss at the old weight."
With the numbers: , the first row of the table.
Level 3: in Python
w, eta, target = 0.0, 0.25, 3.0
for step in range(3):
# ∂L/∂w for the loss (w - 3)²
slope = 2 * (w - target)
# w ← w - η ∂L/∂w
w = w - eta * slope
print(w) # → 1.5 2.25 2.625
learn_one_weight runs the table above; train runs the loop for a real
network.
Why it matters this loop is identical for a spam filter and for a frontier language model. Only the size of the network, the data, and the loss change. When a training run misbehaves, the fault is in one of these four boxes.
Chapter 2
A neuron: a weighted vote followed by a decision
Think of a judge at a cooking contest. They score taste, presentation and originality, but care about taste most, so they weight it more. They add a baseline ("everyone starts at half a point"), then apply a rule to turn the total into a verdict ("negative totals count as zero"). That is a neuron: weights are how much each input matters, the bias is the baseline, and the activation function is the rule.
Worked example: inputs (2, 3), weights (0.5, 1.0), bias 0.5. Weighted sum: 2 × 0.5 + 3 × 1.0 + 0.5 = 4.5. The ReLU rule keeps positive values, so the output is 4.5.
Figure 2 · Diagram
flowchart LR X1((x1 = 2)) -- "× w1 = 0.5" --> S["Sum + bias<br/>1.0 + 3.0 + 0.5 = 4.5"] X2((x2 = 3)) -- "× w2 = 1.0" --> S B((bias 0.5)) --> S S --> A["Activation φ<br/>ReLU(4.5) = 4.5"] A --> Y((y = 4.5))
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the -th input | , | |
| the weight on the -th input | , | |
| a counter over the inputs | 1, 2 | |
| "add up the following for every " | ||
| the bias (the baseline added to every sum) | 0.5 | |
| the dot product: multiply matching entries of two lists, then add | ||
| "phi", the activation function | ReLU | |
| the neuron's output | 4.5 | |
| a batch of inputs, one example per row | shape (batch, inputs) | |
| a layer's weights, one column per neuron | shape (inputs, neurons) | |
| matrix multiply: every row of dotted with every column of | shape (batch, neurons) | |
| every neuron's output for every example | shape (batch, neurons) |
In words: "multiply each input by its weight, add them up, add the bias, and pass the total through the activation function."
With the numbers: $y = \text{ReLU}(0.5 \cdot 2 + 1.0 \cdot 3 + 0.5) = \text{ReLU}(4.5) = 4.5$.
Level 3: in Python
x = [2, 3]
w = [0.5, 1.0]
b = 0.5
# ReLU: keep positives, zero the rest
def phi(z): return max(0.0, z)
# w · x = Σ_i w_i x_i
sum(w_i * x_i for w_i, x_i in zip(w, x)) # → 4.0
# y = φ(w · x + b)
phi(sum(w_i * x_i for w_i, x_i in zip(w, x)) + b) # → 4.5
Why it matters every dense layer in every model is this, including the
feed-forward half of each transformer block, where most of a language
model's parameters live. MLP.forward is two of these layers in six lines.
Chapter 3
Backpropagation: passing the blame backwards
Picture an assembly line that produced a faulty product. The inspector at the end measures how bad it is and passes the complaint back. Each station works out how much of the fault was its own doing, based on what it received and what it did to it, and passes the rest further back. Every station ends up knowing exactly how to adjust its own machine. Backpropagation is that blame-passing, done with derivatives.
Worked example, continuing the neuron: the target was 5, so the loss is (5 − 4.5)² = 0.25.
| link | local derivative | running blame |
|---|---|---|
| loss → y | 2(y − 5) = −1 | −1 |
| y → z (ReLU, z > 0) | 1 | −1 |
| z → w1 | x1 = 2 | −2 |
| z → w2 | x2 = 3 | −3 |
| z → b | 1 | −1 |
One step with learning rate 0.01 moves weights to (0.52, 1.03) and bias to 0.51. The new output is 4.64 and the loss drops from 0.25 to 0.1296. The weight whose input was bigger (3) received the bigger share of blame and moved more.
Figure 3 · Diagram
flowchart RL L["Loss (5 − y)² = 0.25"] -- "dL/dy = −1" --> Y[y = ReLU z] Y -- "dy/dz = 1" --> Z[z = w·x + b] Z -- "× x1 = 2 → dL/dw1 = −2" --> W1[w1] Z -- "× x2 = 3 → dL/dw2 = −3" --> W2[w2] Z -- "× 1 → dL/db = −1" --> Bb[b]
Level 3: the formula and its symbols
This is the chain rule: when a change passes through several steps, the overall rate of change is the product of each step's rate of change. If turning a knob moves a gear 3× as fast, and that gear moves a needle 2× as fast, the knob moves the needle 6× as fast.
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the loss | 0.25 | |
| the target (the right answer) | 5 | |
| the neuron's output | 4.5 | |
| the weighted sum before the activation | 4.5 | |
| how the loss changes as the output changes | ||
| how the output changes as the sum changes: the activation's slope | ReLU slope at 4.5 = 1 | |
| how the sum changes as weight changes: just its input | ||
| the activation's derivative (slope) | 1 for ReLU when |
In words: "the blame on a weight is how much the loss cares about the output, times how much the output cares about the sum, times how much the sum cares about that weight."
With the numbers: for : ; for : .
In Python:
x, t, z = [2, 3], 5, 4.5
# ReLU(4.5)
y = max(0.0, z)
# ∂L/∂y
dL_dy = 2 * (y - t)
# φ'(z): ReLU's slope
dy_dz = 1 if z > 0 else 0
# × ∂z/∂w_i = x_i, for w_1 and w_2
[dL_dy * dy_dz * x_i for x_i in x] # → [-2.0, -3.0]
# one step, learning rate 0.01
w = [0.5 - 0.01 * -2.0, 1.0 - 0.01 * -3.0]
b = 0.5 - 0.01 * -1.0
y_new = w[0] * x[0] + w[1] * x[1] + b
# the new output and the smaller loss
round(y_new, 2), round((t - y_new) ** 2, 4) # → (4.64, 0.1296)
one_neuron_worked_example computes every row of the table.
Why it matters PyTorch's loss.backward() does exactly this, for
billions of weights, by recording the forward computation and replaying it
backwards. Knowing it is just the chain rule is what lets you reason about
vanishing gradients, memory use, and why some architectures train and others
don't.
Chapter 4
Why the nonlinearity is essential
Stack three sheets of tinted glass and you get one darker sheet of glass: no arrangement of flat panes can bend light around a corner. Linear layers are flat panes. However many you stack, the result is one linear layer, which can only separate classes with a straight line. The activation function is the bend.
Worked example in one dimension: a layer that multiplies by 3, followed by a layer that multiplies by 2, is the same as one layer that multiplies by 6. With a ReLU between them, inputs −1 and 1 give 0 and 6. No single multiplication does that.
Figure 5 · Diagram
flowchart LR
subgraph NoAct["Without activation"]
a1[X] --> a2[·W1] --> a3[·W2] --> a4[·W3] --> a5["= X·(W1W2W3)<br/>one matrix"]
end
subgraph WithAct["With activation"]
b1[X] --> b2[·W1] --> b3[φ] --> b4[·W2] --> b5[φ] --> b6[·W3] --> b7[curved<br/>decision boundary]
end
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the inputs | or | |
| two layers' weights | 3 and 2 | |
| the two layers multiplied into one | 6 | |
| an activation between the layers | ReLU | |
| any single weight you might try instead | none works | |
| "is not equal to" |
In words: "two linear layers in a row are the same as one linear layer, but put an activation between them and no single layer can copy them."
With the numbers: without ReLU, and , exactly . With ReLU, and . A single weight would need and at once: impossible.
Level 3: in Python
W_1, W_2 = 3, 2
# ReLU
def phi(z): return max(0, z)
# (X W_1) W_2 ...
[(x * W_1) * W_2 for x in [-1, 1]] # → [-6, 6]
# ... equals X (W_1 W_2): one layer of 6
[x * (W_1 * W_2) for x in [-1, 1]] # → [-6, 6]
# φ(X W_1) W_2: no single W' gives 0 and 6
[phi(x * W_1) * W_2 for x in [-1, 1]] # → [0, 6]
Figure 4 · Drawn from the lesson's code
Two-moons decision boundaries: logistic regression's straight line misclassifies the moon tips (88%), while one hidden tanh layer bends around the gap (99.75%)
In code: make_moons builds the two interleaving half-moons (and make_xor the four-point XOR puzzle); train_logistic_regression fits the straight-line baseline in the left panel.
Why it matters "depth" is only worth anything because of the
activations between layers. linear_stack_collapses shows five stacked
linear layers matching their single-matrix product to 1e-16.
Chapter 5
Activation functions: the rule after the sum
Think of different kinds of switch. ReLU is a one-way valve: water flows freely forward, not at all backward. Sigmoid is a dimmer that squeezes any input into 0–1 but barely moves at the extremes. Tanh is the same dimmer centred on zero. GELU is a valve that leaks a little when nearly closed.
Worked example, values you can check on a calculator (the ′ mark, as in sigmoid′, means "the slope of"):
| z | ReLU | sigmoid | sigmoid' | tanh' | GELU |
|---|---|---|---|---|---|
| −1 | 0 | 0.269 | 0.197 | 0.420 | −0.159 |
| 0 | 0 | 0.5 | 0.25 (its maximum) | 1 (its maximum) | 0 |
| 1 | 1 | 0.731 | 0.197 | 0.420 | 0.841 |
Figure 6 · Drawn from the lesson's code
Activations and their slopes: sigmoid's slope peaks at 0.25 and sigmoid and tanh slopes vanish past |z| = 3, while ReLU's slope is a clean step from 0 to 1
| Function | Formula | Range | Where it's used |
|---|---|---|---|
| ReLU | max(0, z) | [0, ∞) | Hidden layers of CNNs and MLPs. Fast. A neuron stuck negative gets zero gradient forever ("dying ReLU"). |
| GELU | z·Φ(z), Φ = normal CDF | ≈[−0.17, ∞) | Transformers (BERT, GPT-2 and most since). |
| Sigmoid | 1 / (1 + e^−z) | (0, 1) | Binary outputs, and the gates inside LSTMs. |
| Tanh | (e^z − e^−z)/(e^z + e^−z) | (−1, 1) | RNN hidden states; zero-centred. |
| Softmax | e^{z_i} / Σ e^{z_j} | probabilities | The output layer over classes or tokens, and inside attention. |
Reading the formulas: e is Euler's number, about 2.718, and e^z means
"2.718 raised to the power z", which is always positive and grows fast.
Φ(z) is the fraction of a bell curve (the standard normal distribution) that
lies below z: Φ(0) = 0.5, Φ(1) = 0.841. Softmax turns a list of scores
into shares that are all positive and add up to 1: raise e to each score,
then divide each by the total (Σ means "add them all up"). See
primer.notation for each symbol, and primer.ml.attention for softmax
worked through by hand.
In code: relu, sigmoid, tanh and gelu compute each rule, and relu_grad, sigmoid_grad, tanh_grad and gelu_grad compute its slope; gelu_tanh is the cheaper approximation GPT-2 uses, and softmax subtracts the largest score first so nothing overflows.
Why it matters activation choice decides whether gradients survive a
deep stack. Sigmoid everywhere is why deep networks were hard to train
before 2010; ReLU and residual connections are why they aren't now (see
primer.ml.deep_nets).
Chapter 6
Backprop through two layers
Now the assembly line has two stations. The blame from the end first reaches the output weights, then travels back through the hidden layer to the first weights. On the way it passes through the hidden activation, which can shrink it, just as a station that barely changed the product can take only a little of the blame.
Worked example, one input, one hidden unit, one output: x = 1, w1 = 0.5, tanh, w2 = 2, sigmoid, target 1.
| step | value |
|---|---|
| z1 = x·w1 | 0.5 |
| h = tanh(0.5) | 0.4621 |
| z2 = h·w2 | 0.9242 |
| ŷ = σ(0.9242) | 0.7159 |
| dL/dz2 = ŷ − y | −0.2841 |
| dL/dw2 = h · dL/dz2 | −0.1313 |
| dL/dh = dL/dz2 · w2 | −0.5682 |
| dL/dz1 = dL/dh · (1 − h²) | −0.5682 × 0.7864 = −0.4469 |
| dL/dw1 = x · dL/dz1 | −0.4469 |
Figure 7 · Diagram
flowchart LR
subgraph Forward
X[X] --> Z1["z1 = X·W1 + b1"] --> H["h = tanh z1"] --> Z2["z2 = h·W2 + b2"] --> P["ŷ = σ z2"] --> L["L = BCE ŷ, y"]
end
L -. "ŷ − y" .-> dZ2[dL/dz2]
dZ2 -. "hᵀ · dz2" .-> dW2[dL/dW2]
dZ2 -. "dz2 · W2ᵀ" .-> dH[dL/dh]
dH -. "⊙ 1 − h²" .-> dZ1[dL/dz1]
dZ1 -. "Xᵀ · dz1" .-> dW1[dL/dW1]
H -. cached .-> dW2
H -. cached .-> dZ1
X -. cached .-> dW1
MLP.backward, and the worked table
above is that same path with numbers. Notice the three "cached" arrows: the
backward pass reuses h and X from the forward pass. That dependency is why
training must keep every layer's activations in memory, and inference does
not.For a batch, with binary cross-entropy (the loss for yes/no predictions,
; see primer.ml.losses), which
fuses with sigmoid into the simple error ŷ − y:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example (one example, one unit) |
|---|---|---|
| inputs, one row per example | ||
| first and second layer weights | 0.5, 2 | |
| each layer's weighted sum, before its activation | 0.5, 0.9242 | |
| hidden activations, | 0.4621 | |
| "y-hat", the prediction | 0.7159 | |
| the target | 1 | |
| sigmoid, squashes any number into 0 to 1 | ||
| hyperbolic tangent, squashes into −1 to 1 | ||
| "transpose": flip a matrix so rows become columns, so the shapes line up for the multiply | a scalar is its own transpose | |
| multiply element by element (not a matrix multiply) | ||
| the slope of tanh at | 0.7864 |
In words: "the error at the output is prediction minus target; each layer's weight gradient is that layer's input times the error arriving at it; to send the error one layer further back, multiply by the weights it came through and by the activation's slope."
With the numbers: ; ; ; ; .
Level 3: in Python
import math
X, W_1, W_2, y = 1, 0.5, 2, 1
z_1 = X * W_1
h = math.tanh(z_1)
z_2 = h * W_2
# σ(z_2)
y_hat = 1 / (1 + math.exp(-z_2))
# ŷ - y
dL_dz2 = y_hat - y
# hᵀ ∂L/∂z_2
dL_dW2 = h * dL_dz2
# ∂L/∂z_2 W_2ᵀ
dL_dh = dL_dz2 * W_2
# ⊙ (1 - h²), tanh's slope
dL_dz1 = dL_dh * (1 - h ** 2)
# Xᵀ ∂L/∂z_1
dL_dW1 = X * dL_dz1
[round(v, 4) for v in (dL_dz2, dL_dW2, dL_dh, dL_dz1, dL_dW1)] # → [-0.2841, -0.1313, -0.5682, -0.4469, -0.4469]
The hand-written gradients are checked two ways: a numerical gradient
check (nudge each weight by ±ε and measure the loss change; see
gradient_check) and, in the tests, PyTorch autograd.
In code: tiny_two_layer_example computes every row of the worked table; MLP holds both layers' weights and biases, MLP.loss scores a batch with binary_cross_entropy, and numerical_gradient is the slow nudge-every-weight answer that gradient_check compares against.
Why it matters the factor (1 − h²) is where gradients shrink. Chain fifty of them and the first layers hear almost nothing: the vanishing gradient problem. The cached activations are why training a model needs several times the memory of serving it.
Chapter 7
Batch, step, epoch
Imagine studying a deck of 400 flashcards. You work through a pile of 32, then pause to update your notes; that pause is a step. The pile is a batch. Going through the whole deck once is an epoch; then you shuffle and start again.
Worked example: 400 examples in batches of 32 is 12 full batches plus one of 16, so ⌈400/32⌉ = 13 steps per epoch. Two epochs: 26 steps.
Figure 9 · Diagram
flowchart LR E[Epoch: one pass over all 400 examples] --> SH[Shuffle] SH --> B1[Batch 1<br/>32 examples] --> S1[Step 1<br/>update weights] S1 --> B2[Batch 2] --> S2[Step 2] S2 --> D[...] --> B13[Batch 13<br/>last 16 examples] --> S13[Step 13] S13 -->|next epoch| E
Figure 8 · Drawn from the lesson's code
Training loss per epoch on a log axis: slow at first, a steep fall once hidden units find features, then a flat tail near zero
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| number of training examples | 400 | |
| batch size | 32 | |
| "ceiling": round up to the next whole number (the last, smaller batch still counts as a step) |
In words: "an epoch takes as many steps as it takes batches of size B to cover N examples, rounding up."
With the numbers: steps per epoch; 2 epochs = 26 steps.
Level 3: in Python
import math
N, B, epochs = 400, 32, 2
N / B # → 12.5
# ⌈N / B⌉: the last, smaller batch still counts
math.ceil(N / B) # → 13
epochs * math.ceil(N / B) # → 26
In code: train reshuffles the data every epoch, takes one plain gradient step per batch and records the loss after each epoch; MLP.accuracy reports the fraction of examples classified correctly.
Why it matters larger batches give smoother gradient estimates and keep GPUs busy, but need more memory. Pretraining a language model runs roughly one epoch over trillions of tokens (it rarely sees the same text twice); fine-tuning runs several epochs over a small dataset, which is why fine-tunes can overfit.
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Why can't you just stack linear layers?Think it through, then reveal
Their composition is a single linear map (W1·W2 is just another matrix), so extra layers add parameters but no expressive power.
Question 2What does backprop actually compute?Think it through, then reveal
The gradient: for every weight, how much a tiny change would change the loss. It applies the chain rule from the output backward, reusing values cached in the forward pass.
Question 3Why is ReLU preferred over sigmoid in hidden layers?Think it through, then reveal
Sigmoid's derivative is at most 0.25 and near 0 when saturated, so gradients shrink multiplicatively with depth. ReLU's derivative is exactly 1 for positive inputs. The cost: neurons whose input stays negative get zero gradient and "die".
Question 4Why does training need more memory than inference?Think it through, then reveal
The backward pass needs every layer's activations from the forward pass, so they must be kept until the gradients are computed. Inference can discard each activation as soon as the next layer has used it.
Question 5What's the gradient of sigmoid + binary cross-entropy with respect to the logit?Think it through, then reveal
ŷ − y: prediction minus target. Clean, bounded and cheap, which is why the two are always fused.
Question 6Batch vs. step vs. epoch?Think it through, then reveal
A batch is the examples per update, a step is one update, an epoch is one full pass over the data.
Primary sources
The papers behind this lesson
Rumelhart, Hinton & Williams, Learning representations by back-propagating errors (Nature, 1986): Showed that the chain rule, run backwards through a multi-layer network, trains hidden units to discover useful internal features.
The paper ↗Hendrycks & Gimpel, Gaussian Error Linear Units (GELUs) (2016): Introduced GELU, the smooth ReLU that became the default activation in transformers.
The paper ↗Researcher's shelf
Further reading
- CS231n notes, Backpropagation, intuitions: https://cs231n.github.io/optimization-2/
- CS231n notes, Neural Networks Part 1: https://cs231n.github.io/neural-networks-1/
- Andrej Karpathy, micrograd (autograd in about 100 lines): https://github.com/karpathy/micrograd
- Andrej Karpathy, The spelled-out intro to neural networks and backpropagation: https://www.youtube.com/watch?v=VMj-3S1tku0
- 3Blue1Brown, Neural networks series: https://www.3blue1brown.com/topics/neural-networks
- Michael Nielsen, Neural Networks and Deep Learning, ch. 2: http://neuralnetworksanddeeplearning.com/chap2.html
- Hendrycks & Gimpel, GELU (2016): https://arxiv.org/abs/1606.08415
- PyTorch autograd tutorial: https://pytorch.org/tutorials/beginner/blitz/autograd_tutorial.html
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.