The lesson in one minute
What you'll be able to explain
- Overfitting: training error falls while validation error rises; the model memorised noise. Underfitting: both errors are high.
- Fixes for overfitting: more data, early stopping, dropout, weight decay (L2), L1 for sparsity, data augmentation.
- Bias is systematic error (too simple); variance is sensitivity to the training sample (too flexible).
- Train fits weights, validation tunes choices, test is touched once; cross-validation for small data.
- Leakage (duplicates across splits, features from the future) gives great offline scores that collapse in production.
Level 1
The practitioner's guide
In one sentence
Regularization is everything you do to make a model learn the pattern in its training data rather than memorise the examples, and the way you know it is working is the gap between the error on data the model trained on and the error on data it has never seen.
When you need it
Whenever you train or fine-tune anything on less data than the model could memorise, which in practice is always: a fine-tune on a few thousand examples, a classifier on a table of ten thousand rows, a forecasting model on a year of history. The tell is a training error that keeps falling while the validation error turns and climbs. This lesson's polynomial sweep shows the whole story in one table: at degree 5 the training error is 0.013 and the validation error 0.057; at degree 9 the training error has dropped to 0.004 and the validation error has risen twelve-fold to 0.71; at degree 11 the training error is zero and the validation error is 7,631. Its over-sized network bottoms out on validation near epoch 300 and gets steadily worse for the next 1,700 epochs while its training loss keeps improving. The other tell is the opposite one, and it hides in a good offline score: a model that looked superb in evaluation and collapsed in production has usually seen the answers, and that is leakage, which the same lesson covers. You do not need more regularization when both errors are high; that is underfitting, and it wants more capacity, more features or more training, the opposite fix.
Your options
From the cheapest to the most involved:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| A held-out validation set (and a test set touched once) | Measures the gap that every other option is judged by | An honest estimate, as long as the test set is not used to choose | 15% to 30% of your data, and discipline | Your data pipeline |
| Early stopping | Watches validation loss, stops after patience bad checks, restores the best weights |
You ship the best model seen, not the last | A validation pass per epoch; nothing in the model | The training loop: a callback in every framework |
| More data, or augmented data | Averages the noise away | Tames variance in flexible models; does nothing for a model too simple to fit the pattern | Collection or labelling time | Your data |
| L2 penalty (weight decay) | Taxes the square of every weight, so all of them shrink a little | Smaller, smoother weights; nothing is eliminated | One hyperparameter, λ | The optimiser: on by default in AdamW |
| Dropout | Zeroes each activation with probability p on every training step, scaled so evaluation is a plain pass-through | No neuron becomes indispensable | Slower convergence; a bug if left on at evaluation | The model, between layers |
| L1 penalty (lasso) | Charges a flat fee per unit of weight, so small weights go to exactly zero | A sparse model that names its features | Harder optimisation; one hyperparameter | Linear and tabular models, feature selection |
| Less capacity | A smaller model, a lower polynomial degree, fewer features | Less to memorise with | Possibly too little to learn the pattern | The architecture |
| Cross-validation | Trains k models, each with one fold held out, and averages their scores | A steadier estimate than one split when data is small | k training runs | Small tabular problems; scikit-learn's cross-validation guide |
| Leak-proof splitting | Deduplicates before splitting, splits by time or by user, drops features unknown at prediction time | An offline score that survives production | Thinking about the timeline of every feature | Your data pipeline, before any training |
How to choose
Diagnose first, because the two failures want opposite cures.
- Both errors high: underfitting. Add capacity, features or training time. Regularizing further makes it worse.
- Training low, validation high: overfitting. Early stopping first, because it costs nothing; then weight decay (which you probably already have) and dropout; then more data if you can get it.
- A tabular model you must explain, or hundreds of features of which a few matter: L1, which in this lesson zeroes 5 of 7 useless weights while L2 zeroes none.
- Fewer than a few thousand examples: cross-validate rather than trust one split.
- A model that was wonderful offline: before celebrating, check for twins across the split and for columns filled in after the outcome.
- Whatever you pick, the test set is touched once, at the end. Every
decision made by looking at it turns it into a second validation set, and
the benchmark version of this failure is Goodhart's law in
primer.ml.benchmarks.
What it costs
Early stopping costs one validation pass per epoch and stops training sooner, so it usually saves compute. Weight decay and L1 cost nothing at inference and one hyperparameter each; PyTorch's AdamW defaults weight decay to 0.01. Dropout slows training, because each step trains a thinned network, and the original transformer used p = 0.1 rather than anything heavier. Cross-validation multiplies training cost by the number of folds. The dear cost is data: the validation and test sets are examples you cannot train on, and with 100 examples a 70 / 15 / 15 split leaves 15 for each. Leakage costs the most of all, late: this lesson's fraud model with a future feature scores 98% offline and 49% in production, worse than the honest model's 61%.
What breaks
- Dropout left on at evaluation. Predictions become noisy and change
from call to call. Switch the model to evaluation mode; the lesson's
dropoutpasses activations through unchanged there. - Early stopping with no restore. Stopping is half the job; the weights at the stop are the ones that just failed twice. Keep the snapshot from the best check.
- Patience too short. A noisy validation curve stops training on a blip. Smooth it, or raise the patience.
- Tuning against the test set. Each look fits the model to it a little. Choose with validation, report with test.
- Duplicates across the split. A nearest-neighbour memoriser scores 87% on data whose labels are 30% noise, when no honest model can beat about 70%; deduplicated, it scores 50%. Deduplicate before splitting.
- A feature from the future.
chargeback_filedis recorded after the outcome. Ask of every column: would I know this at prediction time? - A random split on time-ordered data. Predicting next month from a shuffle of all months is a leak. Split by time when the task is the future, by user when the task is new users.
- Regularizing an underfit model. It cannot get better by learning less. Look at both curves before reaching for the penalty.
In the wild
Weight decay is on by default in the optimisers that train
transformers: torch.optim.AdamW takes weight_decay with a default of
0.01, and primer.ml.optimizers shows why decoupling it from the gradient
step mattered. Early stopping is a callback everywhere it is offered;
Keras's EarlyStopping has patience and restore_best_weights, which
default to 0 and False, so a naive call gets neither. Dropout (Srivastava
et al., 2014) trains at p = 0.1 in the original transformer and is switched
off by model.eval() in PyTorch. The lasso is Tibshirani (1996), and
scikit-learn's cross-validation guide is the standard reference for folds
that respect groups and time. Kapoor and Narayanan (2022) catalogued the
kinds of leakage and how widely they inflate published results, and Belkin
et al. (2018) documented double descent, where very over-parameterised
models generalise well again, which is why the modern recipe is a large
model plus regularization rather than a small model.
Go deeper
Level 2 fits the memorising cubic by hand, sweeps polynomial degree to draw the valley in validation error, replays early stopping on a list of losses, splits error into bias and variance over 300 training sets, builds dropout and the two penalties in a few lines each, cuts data into folds, and reproduces both leaks with a nearest-neighbour memoriser. If you only needed to diagnose a curve and pick a fix, you are done.
Level 2
How it works, from scratch
Two students prepare for an exam. One learns the ideas. The other memorises last year's paper, answer by answer, and scores 100% on it in practice. On the real exam, with new questions, the first student does fine and the second falls apart. A model that memorises its training examples, noise and all, instead of learning the pattern behind them is overfitting. It looks brilliant on the data it has seen and fails on data it hasn't. Regularization is everything we do to push a model toward the first student.
Worked example: four training points (0, 0), (1, 1), (2, 0), (3, 1), a zig-zag that is really "about 0.5, plus noise".
| model | error on the 4 training points | prediction at x = 4 |
|---|---|---|
| flat line at 0.5 (the pattern) | 0.25 | 0.5 |
| cubic through all four points (memorised) | 0 | 8 |
The cubic is perfect on the training data and absurd one step beyond it. You can check the 8 by hand with a difference table: the values 0, 1, 0, 1 have differences 1, −1, 1, then −2, 2, then 4; a cubic keeps that last difference constant, so extending the table gives 6, then 7, then 8.
Figure 2 · Diagram
flowchart LR
D["Training data =<br/>pattern + noise"] --> M{Model capacity}
M -->|too little| U["Underfits:<br/>misses the pattern"]
M -->|about right| G["Generalizes:<br/>learns the pattern"]
M -->|too much, unchecked| O["Overfits:<br/>memorises the noise"]
G --> N[Good on new data]
U & O --> B[Bad on new data]
Figure 1 · Drawn from the lesson's code
Three polynomial fits to 12 noisy sine samples: the degree-1 line misses the curve, the cubic follows it, and the degree-11 polynomial hits every dot but swings wildly between them
The two numbers that tell these apart are errors on two different sets of data:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| how many training and validation examples | 4 training points | |
| "for every example in the training set" | ||
| the true value | 0, 1, 0, 1 | |
| the model's prediction | 0.5 everywhere for the flat line | |
mean squared error (see primer.ml.losses) |
0.25 for the flat line |
In words: "training error is the average squared miss on the data the model learned from; validation error is the same average on data it has never seen."
With the numbers: flat line: . Cubic: every miss is 0, so training error is 0; its validation error on any point beyond x = 3 is enormous.
Level 3: in Python
# the true values at x = 0, 1, 2, 3
y = [0, 1, 0, 1]
# the flat line predicts 0.5 everywhere
y_hat = [0.5, 0.5, 0.5, 0.5]
n = len(y)
# (1/n) Σ (y_i - ŷ_i)²
sum((y_i - y_hat_i) ** 2 for y_i, y_hat_i in zip(y, y_hat)) / n # → 0.25
In code: zigzag_fit_error and zigzag_predict fit a polynomial of any
degree to the four zig-zag points and report its training error and its
prediction; true_function and make_curve_data draw the noisy sine
samples in the figure.
Why it matters training error alone always rewards memorisation. The gap between training and validation error is the single most useful diagnostic in machine learning.
Chapter 1
Underfitting vs. overfitting: finding the middle
Goldilocks tries three bowls of porridge: too cold, too hot, just right. Model capacity works the same way, and you find "just right" by measuring, not guessing: sweep the capacity and watch both errors.
Worked example: 12 noisy points from a sine wave, 300 fresh points for validation, polynomials of increasing degree.
| degree | training error | validation error | diagnosis |
|---|---|---|---|
| 1 | 0.20 | 0.30 | underfit: both high |
| 3 | 0.021 | 0.065 | about right |
| 5 | 0.013 | 0.057 | best: lowest validation error |
| 9 | 0.004 | 0.71 | overfit: validation 12× worse than at degree 5 |
| 11 | 0.0000 | 7,631 | memorised |
Figure 3 · Drawn from the lesson's code
Error against polynomial degree on a log scale: training error only falls, while validation error is lowest at degree 5 and then climbs steeply
In code: polynomial_errors fits by least squares (choosing the
coefficients that minimise training MSE) and reports both errors.
Why it matters "both errors high" and "training low, validation high" need opposite fixes. Underfitting wants more capacity, more features or more training; overfitting wants more data, regularization or early stopping. Diagnosing which one you have is the point.
Chapter 2
Early stopping: take the cake out before it burns
A baker checks the cake with a toothpick every few minutes and takes it out when it comes out clean. Leave it longer and it burns, however good it was a minute ago. Training a flexible model is similar: validation error falls while the model learns the pattern, then rises once it starts memorising. Early stopping watches validation error during training, stops once it has failed to improve for a few checks (the patience), and keeps the weights from the best check.
Worked example: validation losses per epoch 1.0, 0.8, 0.6, 0.55, 0.58, 0.65, 0.7 with patience 2. Epoch 3 is the best. Epochs 4 and 5 are both worse, so training stops at epoch 5 and the weights from epoch 3 are kept.
Figure 5 · Diagram
flowchart LR
E[End of epoch] --> V[Measure validation loss]
V --> Q{Better than best?}
Q -->|yes| S[Save weights,<br/>reset counter]
Q -->|no| C[Counter + 1]
C --> P{Counter = patience?}
P -->|no| E2[Next epoch]
P -->|yes| R[Stop; restore<br/>best weights]
S --> E2
Figure 4 · Drawn from the lesson's code
Loss curves for an over-sized network: training loss keeps falling while validation loss bottoms out near epoch 300 and then rises
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the epoch number | 0 to 6 | |
| validation loss after epoch | 0.55 at | |
| "the at which the following is smallest" (the position, not the value) | 3 | |
| "t-star", the best epoch so far | 3 | |
| patience | how many non-improving epochs to tolerate | 2 |
In words: "remember the epoch with the lowest validation loss, and stop
once you've gone patience epochs past it without doing better."
With the numbers: ; at , , so stop.
Level 3: in Python
# validation loss after epoch t = 0, 1, 2, ...
L_val = [1.0, 0.8, 0.6, 0.55, 0.58, 0.65, 0.7]
patience = 2
t_star = 0
for t, loss in enumerate(L_val):
if loss < L_val[t_star]:
# a new best epoch: remember it
t_star = t
if t - t_star >= patience:
# patience used up: stop here
break
t_star, t # → (3, 5)
In code: early_stopping replays a run's validation losses and returns
the best epoch and the epoch where training stops; train_flexible_model
trains the over-sized network in the figure and records both losses.
Why it matters it's the cheapest regularizer there is: no change to the model, just a rule for when to stop. It's used almost everywhere a model is trained for multiple epochs, including fine-tuning language models.
Chapter 3
Bias and variance: two ways to miss the target
Picture two archers. The first groups every arrow tightly, but a hand's width to the left of the bullseye: consistently wrong. That's bias. The second's arrows are centred on the bullseye on average but scattered all over the target: inconsistent. That's variance. A model trained on a different sample of data is like another volley of arrows. Simple models behave like the first archer; very flexible ones like the second.
Worked example: fit polynomials to 300 different random training sets of 30 points each, and measure (averaged over the input range):
| degree | bias² (systematic error) | variance (swing between training sets) |
|---|---|---|
| 1 | 0.174 | 0.022 |
| 3 | 0.0035 | 0.015 |
| 9 | 0.0038 | 4.5 |
Figure 6 · Drawn from the lesson's code
Twenty fits from different training sets: degree-1 lines agree but all miss the sine (high bias), degree-9 curves average to the sine but scatter widely (high variance)
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the true pattern | ||
| "f-hat", the model's prediction, which depends on which training set it saw | one thin line | |
| expected value: the average over many random training sets (and noise) | average of the 300 fits | |
| a noisy observation, plus noise | ||
| the noise variance: error no model can remove |
In words: "a model's expected squared error on new data splits into how far its average prediction is from the truth (bias squared), plus how much its predictions scatter around their own average (variance), plus noise nobody can predict."
With the numbers: degree 1: ; degree 3: ; degree 9: . Degree 3 has the best total.
Level 3: in Python
sigma = 0.3
# σ²: the error no model can remove
noise = sigma ** 2
for degree, bias_sq, variance in [(1, 0.174, 0.022), (3, 0.0035, 0.015), (9, 0.0038, 4.5)]:
# bias² + variance + σ²
total = bias_sq + variance + noise
# two significant figures
print(degree, f"{total:.2g}") # → 1 0.29 3 0.11 9 4.6
In code: bias_variance fits one polynomial to each of many random
training sets and splits their error into bias², variance and noise.
Why it matters it names the two failure modes and explains why more data helps flexible models (averaging tames variance) but not rigid ones (more data doesn't fix bias). Very large neural networks complicate the picture ("double descent": past a point, even bigger models generalize better again), but the vocabulary is universal.
Chapter 4
Dropout: nobody gets to be indispensable
A coach who randomly sends half the team home before each practice forces every player to learn every position; no single star can carry the team. Dropout does this to neurons: during training, each activation is set to zero with probability p on every step, so the network can't rely on any single neuron or fragile combination of neurons.
Worked example: four activations (1, 1, 1, 1), drop rate p = 0.5. Suppose the random mask keeps the last two: (0, 0, 1, 1). The survivors are scaled by 1 / (1 − 0.5) = 2, giving (0, 0, 2, 2). The average is still 1, so the next layer sees the same size signal on average. At evaluation time, nothing is dropped and nothing is scaled.
Figure 7 · Diagram
flowchart LR
subgraph Train["Training step"]
a1((1)) --> k1((0))
a2((1)) --> k2((0))
a3((1)) --> k3((2))
a4((1)) --> k4((2))
end
subgraph Eval["Evaluation"]
b1((1)) --> e1((1))
b2((1)) --> e2((1))
b3((1)) --> e3((1))
b4((1)) --> e4((1))
end
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| a layer's activations | (1, 1, 1, 1) | |
| "h-tilde", the activations after dropout | (0, 0, 2, 2) | |
| the random mask of 0s and 1s | (0, 0, 1, 1) | |
| each mask entry is independently 1 with probability , else 0 (a coin flip) | 50/50 | |
| the drop rate | 0.5 | |
| multiply entry by entry |
In words: "flip a biased coin for each activation, zero the ones that lose, and divide the survivors by the keep probability."
With the numbers: .
Level 3: in Python
import random
random.seed(0)
h = [1, 1, 1, 1]
p = 0.5
# m_i ~ Bernoulli(1 - p): a coin flip each
m = [1 if random.random() < 1 - p else 0 for _ in h]
m # → [0, 0, 1, 1]
# (m ⊙ h) / (1 - p)
[m_i * h_i / (1 - p) for m_i, h_i in zip(m, h)] # → [0.0, 0.0, 2.0, 2.0]
In code: dropout draws the mask and scales the survivors during
training, and passes activations through unchanged at evaluation time.
Why it matters dropout was a key ingredient of the deep-learning
revival and is still used in many models (the original transformer used
p = 0.1). Forgetting to switch it off at evaluation (model.eval() in
PyTorch) is a classic bug that makes predictions noisy.
Chapter 5
L1 and L2 penalties: a tax on large weights
Add a tax on the size of the weights to the loss, and the model must decide whether each weight earns its keep. L2 (ridge, "weight decay") taxes the square of each weight: big weights pay a lot, small ones almost nothing, so everything shrinks a little but nothing is eliminated. L1 (lasso) charges a flat rate per unit of size: small weights can't justify the fee, so they're wiped out entirely, to exactly zero.
Worked example: with independent (orthonormal) features, the penalized weights have a closed form. Start from the unpenalized weights (3, 0.5, −2) and use penalty strength 1:
| weight | L2: divide by 1 + 1 | L1: move 1 toward zero, stop at zero |
|---|---|---|
| 3 | 1.5 | 2 |
| 0.5 | 0.25 | 0 |
| −2 | −1 | −1 |
Figure 8 · Drawn from the lesson's code
Fitted weights on ten features where only three matter: L2 leaves small nonzero weights on every useless feature, L1 sets most of them to exactly zero
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| "choose the weights that make the following smallest" | ||
| the features (one row per example) and targets | ||
| the sum of squared prediction errors | ||
| "lambda", the penalty strength | 1 | |
| the L2 norm squared: sum of squared weights | ||
| the L1 norm: sum of absolute weights |
In words: "fit the data, but pay a penalty proportional to the sum of squared weights (L2) or the sum of absolute weights (L1)."
With the numbers: with orthonormal features the solutions are
and
,
the "soft threshold" in soft_threshold.
Level 3: in Python
import math
w = [3, 0.5, -2]
# λ ("lambda" is taken in Python)
lam = 1
# ‖w‖₂², ‖w‖₁
sum(w_i ** 2 for w_i in w), sum(abs(w_i) for w_i in w) # → (13.25, 5.5)
# L2: shrink every weight
[w_i / (1 + lam) for w_i in w] # → [1.5, 0.25, -1.0]
# L1: sign(w) max(|w| - λ, 0)
[math.copysign(max(abs(w_i) - lam, 0), w_i) for w_i in w] # → [2.0, 0.0, -1.0]
In code: penalised_weights gives both closed forms from the table, and
fit_sparse_problem fits the ten-feature problem in the figure (L1 by
alternating gradient steps with soft_threshold).
Why it matters L2 (as weight decay; see AdamW in
primer.ml.optimizers) is on by default when training transformers. L1
is the tool when you want a sparse, interpretable model that uses only a
few features.
Chapter 6
Train, validation and test: practice, mock and real exams
A student has practice questions (to learn from), a mock exam (to check progress and decide what to revise), and the real exam, sat once. Data is split the same way. Training data fits the weights. Validation data tunes choices like model size, learning rate and when to stop. Test data is touched once, at the end, to estimate real-world performance honestly. If you keep tuning against the test set, it quietly becomes a second validation set and stops being honest.
Worked example: 100 examples split 70 / 15 / 15, shuffled first, with no example in two sets. With only 10 examples, 5-fold cross-validation cuts them into 5 folds of 2; each fold takes a turn as the validation set while the model trains on the other 8, so every example is held out exactly once.
Figure 9 · Diagram
flowchart TB
subgraph K["5-fold cross-validation"]
f1["Run 1: [VAL] train train train train"]
f2["Run 2: train [VAL] train train train"]
f3["Run 3: train train [VAL] train train"]
f4["Run 4: train train train [VAL] train"]
f5["Run 5: train train train train [VAL]"]
end
K --> A[Average the 5 validation scores]
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| number of folds | 5 | |
| which fold is held out | 1 to 5 | |
| score(model, fold) | e.g. accuracy of that model on that fold |
In words: "train k models, each with one fold held out, and average their scores on the fold each one didn't see."
With the numbers: 10 examples, : five models, each trained on 8 and scored on 2; if they score 1.0, 0.5, 1.0, 1.0, 0.5, the CV score is 0.8.
Level 3: in Python
examples = list(range(10))
k = 5
# 5 folds of 2
folds = [examples[2 * i: 2 * i + 2] for i in range(k)]
# each model trains on the other 8
[len(examples) - len(fold) for fold in folds] # → [8, 8, 8, 8, 8]
# score of the model trained without fold i, on fold i
scores = [1.0, 0.5, 1.0, 1.0, 0.5]
# (1/k) Σ_i score_i
sum(scores) / k # → 0.8
In code: train_val_test_split shuffles and cuts the indices into three
disjoint sets, and k_fold yields the training and validation indices for
each fold in turn.
Why it matters with small datasets a single split is noisy, and cross-validation gives a more reliable estimate. With large datasets (or expensive models) one fixed validation set is enough.
Chapter 7
Data leakage: the model saw the answers
A student who glimpsed the answer key scores perfectly on the practice exam and learns nothing. Data leakage is when information that won't be available at prediction time, or information from the test set, sneaks into training. Offline scores look superb; production scores collapse.
Worked examples:
- Duplicates across the split. Every row stored twice, labels 30% noisy (so no honest model can beat about 70%). Split before removing duplicates, and a model that just copies its nearest training row scores 87%: it's reading the answers off each test row's twin. Deduplicate first and the same model scores about 50%.
- A feature from the future. A fraud model is given
chargeback_filed, a column filled in only after a customer disputes a charge, which is to say after the outcome. Offline accuracy: 98%. In production that column is always empty at decision time, and accuracy falls to 49%, worse than an honest model using only the legitimate features (61%).
Figure 10 · Diagram
flowchart LR
subgraph T["Timeline of one transaction"]
direction LR
A[Transaction happens] --> P["Model must decide<br/>(prediction time)"]
P --> O[Fraud confirmed?]
O --> C[chargeback_filed recorded]
end
C -. "leaks backwards into<br/>the training table" .-> P
In code: duplicate_leakage_demo and future_feature_demo run both
examples; the "model" in the first is a one-nearest-neighbour lookup (copy the label
of the closest training row), the purest memoriser there is.
Why it matters leakage is the most common reason a model that looked great in evaluation fails in production. For language-model benchmarks the big version is contamination: test questions that appeared in the pretraining data. The defences are procedural: deduplicate before splitting, split by time or by user when the real task is predicting the future or new users, and ask of every feature "would I actually know this at prediction time?"
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Training loss drops while validation loss rises. What's happening, and what helps?Think it through, then reveal
Overfitting: the model is memorising training-set noise. Stop early (keep the best-validation weights), add regularization (dropout, weight decay), get more or more varied data, or reduce capacity.
Question 2How do you tell underfitting from overfitting?Think it through, then reveal
Underfitting: training and validation errors are both high. Overfitting: training error is low and validation error is much higher. They need opposite fixes.
Question 3Why does dropout scale the surviving activations by 1/(1 − p)?Think it through, then reveal
So the expected value of each activation is the same during training as at evaluation, where nothing is dropped. The evaluation-time network can then be used as is.
Question 4Why does L1 produce exact zeros but L2 doesn't?Think it through, then reveal
L1's penalty has a constant slope, so small weights feel a fixed pull toward zero that the data can't outweigh; they're thresholded to zero. L2's pull is proportional to the weight, so it fades as a weight shrinks and never quite reaches zero.
Question 5Why must the test set be touched only once?Think it through, then reveal
Every decision made by looking at test scores fits the model to the test set a little. Do it repeatedly and the test score stops being an unbiased estimate of performance on new data.
Question 6Give two examples of data leakage.Think it through, then reveal
The same (or near-duplicate) documents in both the training and test splits; and a feature that's only known after the outcome (a chargeback flag in a fraud model). For LLM evaluation: benchmark questions present in the pretraining data.
Primary sources
The papers behind this lesson
Srivastava, Hinton, Krizhevsky, Sutskever & Salakhutdinov, Dropout: A Simple Way to Prevent Neural Networks from Overfitting (JMLR, 2014): Introduced dropout and showed it as a cheap approximation to averaging many networks.
Read the annotated companion →The paper ↗Tibshirani, Regression Shrinkage and Selection via the Lasso (JRSS B, 1996): Introduced the L1 penalty and its property of setting coefficients exactly to zero.
The paper ↗Geman, Bienenstock & Doursat, Neural Networks and the Bias/Variance Dilemma (Neural Computation, 1992): Framed generalization in neural networks as the bias-variance trade-off.
The paper ↗Belkin, Hsu, Ma & Mandal, Reconciling modern machine learning practice and the bias-variance trade-off (2018): Documented "double descent", where very over-parameterized models generalize better again.
The paper ↗Kapoor & Narayanan, Leakage and the Reproducibility Crisis in ML-based Science (2022): Catalogued the kinds of data leakage and how widely they inflate published results.
The paper ↗Researcher's shelf
Further reading
- Goodfellow, Bengio & Courville, Deep Learning, ch. 7 (regularization): https://www.deeplearningbook.org/contents/regularization.html
- CS231n notes, Neural Networks Part 2 (regularization and dropout): https://cs231n.github.io/neural-networks-2/
- Hastie, Tibshirani & Friedman, The Elements of Statistical Learning (free PDF; ch. 3 and 7): https://hastie.su.domains/ElemStatLearn/
- scikit-learn user guide, Cross-validation: https://scikit-learn.org/stable/modules/cross_validation.html
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.