The lesson in one minute
What you'll be able to explain
- Features are directions in activation space, not individual neurons; a dot product with the right direction reads a feature out.
- Probes are small linear classifiers trained on frozen activations. They show information is present, not that the model uses it; score them on unseen data and against a control task.
- The logit lens applies the model's own output layer to intermediate residual states, showing what it would predict after each layer.
- Activation patching copies one activation from a clean run into a corrupted run and measures how much of the right answer returns: the basic causal experiment.
- Superposition: when features are sparse, models store more features than they have neurons, as nearly perpendicular directions, which makes neurons polysemantic.
- Sparse autoencoders learn an overcomplete dictionary with an L1 penalty so each input uses a few latents, recovering features from superposition, at the cost of some unexplained variance.
Level 1
The practitioner's guide
In one sentence
Interpretability is the set of tools that read and change a model's internal activations to find out what it represents and which of those representations cause its answer, as opposed to evaluations, which only measure what it says.
When you need it
Behavioural testing answers "what does the model do?", and for most shipping decisions that is enough. You need to open the hood when the question is "why?": a failure you cannot reproduce from the outside, a model that may be right for the wrong reason (the classifier that detects wolves by the snow behind them), a claim that the model is relying on a protected attribute, or a feature you want to turn up or down without retraining. The tell is that you are about to explain a model's behaviour from its outputs alone and cannot tell two explanations apart. This lesson's toy shows why that is dangerous: a probe reads "tense" out of the hidden state at 98% accuracy on unseen examples, and flipping tense moves the model's output by exactly 0.00, while flipping sentiment moves it by 2.00. What is readable inside a model and what the model uses are different questions, and only an intervention answers the second. One practical limit before anything else: every tool here needs the activations of a model you run yourself. A hosted API gives you tokens, not hidden states, so for a model behind an API your instrument is still the evaluation.
Your options
From the cheapest look to the strongest evidence:
| Option | What it does | What it tells you | What it costs | Where it lives |
|---|---|---|---|---|
| Behavioural evaluation | Runs the model on your tasks and scores the outputs | What the model does, at scale, for any model including hosted ones | A golden set and a scoring rule | primer.agents.evals |
| Linear probe | Trains a tiny classifier on frozen hidden states to read one property | That the property is linearly present at that layer, if it beats a control task on unseen data | Labelled examples (100 sufficed in this lesson), minutes of training | Your code, on an open-weight model |
| Logit lens (and tuned lens) | Applies the model's own output layer to the residual stream after each layer | When and where a prediction forms inside the model; no training needed | Nothing beyond a forward pass; a small trained translator per layer for the tuned version | Any residual-stream model you can hook |
| Activation patching (causal tracing) | Copies one activation from a clean run into a corrupted run and measures how much of the answer returns | Which sites cause the answer: the minimum for a causal claim | Two prompts that differ in one fact, and one forward pass per site tested | Hooking libraries such as TransformerLens |
| Sparse autoencoder (SAE) features | Learns an overcomplete dictionary so each hidden state is a few interpretable directions | What the model's units of meaning are, in a form you can read and steer | A large training run on activations, a penalty to tune, and 11% of the variance unexplained in this lesson's toy | Released SAE suites for open models, or your own training |
| Feature steering | Adds a found direction to the residual stream at run time | Whether a feature causes what its name suggests, and a lever without retraining | The feature must exist first; too strong a push degrades the output | Research demos on production models |
| Circuit analysis | Patches site by site until every step from input to output is accounted for | A complete mechanism for one behaviour | Weeks of expert time for a small model | Research |
How to choose
Start from the question, and reach for the cheapest tool that answers it.
- "Does the model know X?": a probe, scored on held-out examples against a random-label control. In this lesson the control fits 73% of its training set and scores 50% unseen, which is what a probe that only memorised looks like.
- "When does it decide?": the logit lens. In the landmark model the fact appears at the subject word after layer 1 and reaches the last word only after layer 2.
- "Does this part cause the answer?": patching. Reading is not enough: the lens reads "Paris" at 0.995 at a position where patching restores 0%.
- "What are the features, and can I steer one?": an SAE, followed by a steering experiment to check that the feature means what its label says.
- "Is the model relying on the wrong thing in production?": a probe or SAE feature to find the candidate, then patching or steering to confirm it, then a behavioural evaluation to measure the effect at scale.
- Whatever you pick, place the result on the ladder from correlation (probes, the lens) through intervention (patching, steering) to mechanism (a circuit), and claim only the rung you reached.
What it costs
Probes and the logit lens cost almost nothing: frozen activations, a forward pass, and for a probe a few hundred labelled examples and seconds of training. Patching costs one forward pass per site, so a full map is layers times positions; the lesson's is 3 by 3, a real model's is hundreds by thousands, and it is repeated for every prompt pair you try. Sparse autoencoders are the expensive tool: millions of activations, a dictionary far wider than the layer, and a sparsity penalty whose choice is a trade. In this lesson's sweep, a penalty of 0.3 matches every planted feature above 0.99 and leaves 11% of the variance unexplained; a penalty of 0.01 rebuilds everything and matches the worst feature at only 0.80, so a perfect rebuild score is not evidence that the features are real. Circuit-level explanations cost research time measured in weeks per behaviour. And all of it presumes access: none of these tools runs on a model you only reach through an API.
What breaks
- Decodable is not used. The 98%-readable feature with zero effect on the output. Never claim "the model uses X" from a probe; intervene.
- A probe that is too clever. A deep probe can compute the property itself; a probe with no control task can fit noise. Keep probes linear, score them on unseen data, compare with random labels.
- Reading a site nothing reads. The lens can see an answer at a position no later layer consults. Patching restores 0% there.
- A choice that shapes the answer. Patching results depend on how the corrupted prompt is built; SAE features depend on the dictionary's size and penalty, and a bigger dictionary can split one feature into several.
- Redundancy. Models can partly repair themselves when one component is knocked out, so "patching this restores nothing" does not always mean "this plays no role".
- The unexplained remainder. The variance an SAE does not rebuild is model behaviour no feature describes yet, not noise.
- Labels that are guesses. A feature named "Golden Gate Bridge" is a summary of what makes it fire. The name is checked by steering, not by reading.
- Polysemantic neurons. When features outnumber neurons they cannot line up with them, so a single neuron responds to several unrelated things. Look along directions, not at neurons.
In the wild
TransformerLens exposes the internal activations of thousands of open-weight models and lets you cache, edit and replace them as the model runs, which is the plumbing for the lens, probes and patching. Causal tracing (Meng et al., 2022) located factual recall in middle-layer MLPs at the subject's last token in GPT-style models and edited single facts there; Wang et al. (2022) patched their way to a complete circuit for indirect-object identification in GPT-2 small. The tuned lens (Belrose et al., 2023) fixed the logit lens on early layers. Anthropic's Towards Monosemanticity (2023) trained SAEs on a small transformer, and Scaling Monosemanticity (2024) did it inside Claude 3 Sonnet, where turning up one feature made the model bring up the Golden Gate Bridge in almost every answer. Google DeepMind's Gemma Scope releases trained SAEs for every layer of the Gemma 2 2B and 9B base models, so a practitioner can inspect features without training a dictionary.
Go deeper
Level 2 builds each tool on models small enough to check
by hand: a feature as a direction read back by a dot product, a probe
trained by gradient descent with a control task, the logit lens verified
against the real output on primer.ml.transformer.TinyGPT, activation
patching on a two-layer landmark model with a known circuit, superposition
in a five-features-in-two-neurons toy, and a sparse autoencoder that
recovers the planted features only when its penalty is right. If you only
needed to know what these tools can and cannot tell you about a model in
production, you are done.
Level 2
How it works, from scratch
Everyday picture A car makes a strange noise. You can take it for a test drive and note when the noise happens: that is testing the car's behaviour. Or you can open the hood, put a stethoscope on the engine, and swap parts until the noise stops: that is looking at the mechanism. Test drives tell you what the car does. Only the open hood tells you why.
Every other lesson in this primer treats a trained model as something to
build, train or call. Evaluations (primer.agents.evals) are test drives:
they measure what the model says. Interpretability opens the hood. It
asks what the model's billions of internal numbers represent, and which of
them cause the answer. Three reasons to care:
- Debugging. When a model gets something wrong, you want to know which step failed, the same way you read a stack trace rather than just the error message.
- Trust. A model can be right for the wrong reason. An image classifier that "detects" wolves by looking for snow in the background scores well until it meets a wolf on grass.
- Safety. The output is only part of what a model computes. Some questions (is it relying on a stereotype? does it know more than it says?) can only be answered by reading the computation itself.
A tiny worked example. This lesson asks four questions of one kind of sentence, "the Eiffel Tower is in …" → "Paris", each with its own tool, on toy models small enough to check by hand:
| Question | Tool | What our toy shows |
|---|---|---|
| Is a property stored in this hidden state? | a probe | sentiment, read at 99% on unseen examples |
| What would the model say if it stopped at layer ℓ? | the logit lens | "London", then "Paris", then "Paris" |
| Which activations cause the answer? | activation patching | the subject word early, the last word late |
| What are the model's units of meaning? | superposition and sparse autoencoders | five features packed into two neurons, then recovered |
Figure 1 · Diagram
flowchart LR M["A trained model<br/>(frozen)"] --> A["Hidden activations<br/>at every layer and word"] A --> P["Probe:<br/>is property X in here?"] A --> L["Logit lens:<br/>what would it predict now?"] A --> AP["Activation patching:<br/>does this activation cause the answer?"] A --> S["Sparse autoencoder:<br/>which features make up this state?"] P & L --> R["Reading<br/>(correlation)"] AP --> C["Intervening<br/>(causation)"] S --> U["Units of analysis<br/>(what to read and intervene on)"]
Chapter 1
A feature is a direction
Everyday picture A band with two instruments plays through two speakers. The sound engineer pans the guitar mostly to the right speaker and the piano mostly to the left, but each speaker plays a mix of both. If you unplug one speaker you don't lose "the guitar"; you lose part of everything. The instruments are not the speakers. Each instrument is a setting across all speakers, a direction, and the sound in the room is the sum of every instrument times how loudly it plays.
A model's hidden state works the same way. The speakers are neurons: the individual numbers in a layer's output vector. The instruments are features: the things the model has learned to track, such as "this review is positive" or "this verb is in the past tense". Each feature is stored as a direction in the space of neurons, and the hidden state is the sum of every active feature's direction, scaled by how strongly it is present.
A tiny worked example. Take a hidden layer with two neurons and two features whose directions are "positive" = (0.6, 0.8) and "past tense" = (0.8, −0.6). A sentence that is quite positive (0.9) and a little past-tense (0.4) has the hidden state
0.9 · (0.6, 0.8) + 0.4 · (0.8, −0.6) = (0.54 + 0.32, 0.72 − 0.24) = (0.86, 0.48).
Neuron 1 reads 0.86 and neuron 2 reads 0.48. Neither number is "how
positive" or "how past-tense": each neuron is a blend of both features.
To get the features back, take the dot product (multiply matching
entries and add; see primer.notation) of the hidden state with each
direction:
- positive: 0.6 · 0.86 + 0.8 · 0.48 = 0.516 + 0.384 = 0.9
- past tense: 0.8 · 0.86 − 0.6 · 0.48 = 0.688 − 0.288 = 0.4
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the hidden state: one number per neuron | (0.86, 0.48) | |
| how many features are present | 2 | |
| which feature | 1 = positive, 2 = past tense | |
| how strongly feature is present | , | |
| feature 's direction: a list with one number per neuron, of length 1 | ||
| add up one term per feature | two terms | |
| the amount of feature we read back (the hat means "estimated") | 0.9 | |
| dot product: multiply matching entries, then add |
In words: "the hidden state is the sum of each feature's direction times its strength; to read a feature back, dot the hidden state with that feature's direction."
With the numbers: $h = 0.9 \cdot (0.6, 0.8) + 0.4 \cdot (0.8, -0.6) = (0.86, 0.48)\hat{f}_1 = (0.6, 0.8) \cdot (0.86, 0.48) = 0.9$. The read-back is exact here because the two directions are perpendicular (their dot product is 0.6 · 0.8 + 0.8 · (−0.6) = 0) and each has length 1. Hold on to that condition: the section on superposition is about what happens when it fails.
Level 3: in Python
positive = [0.6, 0.8]
past = [0.8, -0.6]
# h = Σ f_i d_i
h = [0.9 * p + 0.4 * q for p, q in zip(positive, past)]
[round(x, 2) for x in h] # → [0.86, 0.48]
# f̂_i = d_i · h
round(sum(d * x for d, x in zip(positive, h)), 2) # → 0.9
round(sum(d * x for d, x in zip(past, h)), 2) # → 0.4
# the directions are perpendicular: their dot product is 0
round(sum(p * q for p, q in zip(positive, past)), 2) # → 0.0
Figure 2 · Drawn from the lesson's code
Two perpendicular arrows for the positive and past-tense directions, and the hidden state (0.86, 0.48) as their weighted sum, with dotted lines dropping onto each arrow at 0.9 and 0.4
That features are directions, and that most of what a model tracks can be
read with a dot product, is called the linear representation
hypothesis. It is a hypothesis, not a law, but it holds often enough to
power everything below. You have met it before: word-vector arithmetic
such as king − man + woman ≈ queen (primer.ml.embeddings.word2vec) works
because "royalty" and "gender" are directions.
In code: compose_features builds a hidden state from feature amounts and read_features reads them back with dot products.
Chapter 2
Probes: can a straight line read it out?
Everyday picture A doctor can't ask your liver how it is doing, but a blood test can measure a marker that tells them. A probe is a blood test for a hidden state: a tiny classifier, trained by you, that looks at a layer's activations and answers one yes/no question, such as "is this review positive?". The model itself is frozen; only the probe learns.
A tiny worked example. A probe for "positive" on a two-neuron layer has weights w = (1, 0.5) and bias b = 0. On the hidden state h = (2, −1):
- Score it: 1 · 2 + 0.5 · (−1) + 0 = 1.5.
- Squash the score into a probability with the sigmoid σ(z) = 1 / (1 + e^−z), which maps any number into the range 0 to 1: σ(1.5) = 1 / (1 + 0.223) = 0.818.
The probe is 82% sure this hidden state belongs to a positive review. On h = (−1, 1) the score is −1 + 0.5 = −0.5 and σ(−0.5) = 0.378: probably negative.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the frozen model's hidden state for one example | (2, −1) | |
| the probe's weights: one per neuron, learned | (1, 0.5) | |
| the probe's bias: one number, learned | 0 | |
| the probe's raw score (its logit) | 1.5 | |
| the sigmoid: turns any score into a probability between 0 and 1 | ||
| Euler's number, about 2.718 | ||
| the probe's probability that the property is present | 0.818 |
In words: "dot the hidden state with the probe's weights, add the bias, and squash the result into a probability."
With the numbers: $p = \sigma(1 \cdot 2 + 0.5 \cdot (-1) + 0) = \sigma(1.5) = 1 / (1 + 0.223) = 0.818$.
Level 3: in Python
import math
w = [1.0, 0.5]
b = 0.0
h = [2.0, -1.0]
# w · h + b
z = sum(wi * hi for wi, hi in zip(w, h)) + b
z # → 1.5
# σ(z)
round(1 / (1 + math.exp(-z)), 3) # → 0.818
# a second hidden state, (−1, 1), reads as probably negative
round(1 / (1 + math.exp(-(-1 * 1.0 + 1 * 0.5))), 3) # → 0.378
This is exactly logistic regression (primer.ml.neural_net), trained
the usual way: show it examples whose answer you know, measure its
binary cross-entropy (primer.ml.losses), and nudge w and b downhill. The
gradient of that loss with respect to the score is simply (p − y), the gap
between the probe's probability and the true 0/1 label, which makes each
update one line of code.
Figure 4 · Diagram
flowchart LR T["Text with a known label<br/>(positive or negative)"] --> M["Frozen model<br/>(no weights change)"] M --> H["Hidden state h<br/>at the layer under study"] H --> P["Probe: σ(w · h + b)"] P --> Y["Probability the label is 'positive'"] Y --> G["Compare with the true label<br/>(cross-entropy)"] G -. "gradient updates w and b only" .-> P
On our toy model. PlantedModel stores two known features in a
32-neuron hidden layer: sentiment and tense, each as ±1 along its own
random direction, plus random noise of size 0.4 on every neuron. No single
neuron holds either feature. A probe trained on just 100 examples reads
sentiment correctly on 99% of 1,000 examples it never saw.
Why keep probes simple, and check them against a control. A probe can succeed for the wrong reason. Give a probe random labels, a control task with nothing real to find, and it still fits 73% of its 100 training examples, because 33 adjustable numbers can memorize a lot of noise. On unseen examples it scores 50%, a coin flip. Two habits follow: always score a probe on examples it never saw, and compare it with a control task, so that "the probe works" means "the hidden state encodes this" rather than "the probe is clever". This is also why probes are kept linear: a powerful probe (a deep network) could compute the property from raw ingredients by itself, and then its success would tell you about the probe, not the model.
Decodable is not the same as used
Everyday picture A library holds a book nobody ever borrows. Finding it on the shelf proves the library has it, not that it shaped anything any reader did.
A tiny worked example. Our toy model's output reads sentiment and, by construction, gives tense a weight of exactly zero. A probe still reads tense at 98% on unseen examples. Now intervene: take 200 examples, flip only one feature's label, keep everything else fixed, and measure how much the model's output moves.
| Flip | Probe accuracy for it | Output moves by |
|---|---|---|
| sentiment | 99% | 2.00 |
| tense | 98% | 0.00 |
Figure 3 · Drawn from the lesson's code
Left, probe accuracy: sentiment 0.99 and tense 0.98 on unseen examples, random labels 0.73 on training examples but about 0.5 on unseen ones. Right, flipping sentiment moves the output by 2 and flipping tense moves it by 0
A probe finding information does not mean the model uses it. To find out what the model uses you have to change something and watch the output, which is what the rest of this lesson does.
In code: train_probe fits a Probe by gradient descent on frozen hidden states from PlantedModel; probe_report trains the three probes, and flip_effect performs the intervention.
Chapter 3
The logit lens: reading the model's mind mid-thought
Everyday picture A writer keeps every draft of an article. Draft 1 says the landmark is "in a European capital", draft 2 says "probably Paris", the final says "Paris". Reading the drafts shows when the writer made up their mind. The logit lens reads a language model's drafts.
It works because of how a transformer is built (primer.ml.transformer).
Each word has a running vector, the residual stream. Every layer reads
it and adds a correction to it, rather than replacing it. At the very
end, the output layer (the unembedding) turns the final vector into
one score per word in the vocabulary. Because every layer writes into the
same stream, you can apply that final step early, after any layer, and ask
"what would the model predict if it stopped here?"
A tiny worked example. A vocabulary of three words, Paris, London and Rome, a residual stream of two numbers, and an unembedding that scores Paris = first number, London = second number, Rome = minus the first:
| After | Residual | Scores (Paris, London, Rome) | Probabilities | Top guess |
|---|---|---|---|---|
| the embedding | (0.2, 0.3) | (0.2, 0.3, −0.2) | (0.36, 0.40, 0.24) | London |
| layer 1, which adds (0.6, 0) | (0.8, 0.3) | (0.8, 0.3, −0.8) | (0.55, 0.34, 0.11) | Paris |
| layer 2, which adds (1.2, −0.3) | (2.0, 0.0) | (2.0, 0.0, −2.0) | (0.87, 0.12, 0.02) | Paris |
The first draft is a vague "some capital" that happens to lean London; layer 1 tips it to Paris; layer 2 commits.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| which layer we stop after (0 = straight after the embedding) | 2 | |
| the residual stream after layer , for one word | (2.0, 0.0) | |
the model's final normalization (primer.ml.transformer); our 2-number toy has none, so here LN(h) = h |
(2.0, 0.0) | |
| the unembedding: one row per vocabulary word, dotted with the state to give that word's score | rows (1, 0), (0, 1), (−1, 0) | |
| the scores (logits), one per word | (2, 0, −2) | |
| softmax | turns scores into probabilities that add up to 1 (primer.ml.attention) |
(0.87, 0.12, 0.02) |
| what the model would predict if it stopped after layer | Paris at 0.87 |
In words: "take the residual stream partway up, push it through the model's own final normalization and output layer, and read off the probabilities."
With the numbers: after layer 2, $W_U h_2 = (1 \cdot 2 + 0 \cdot 0,; 0 \cdot 2 + 1 \cdot 0,; -1 \cdot 2 + 0 \cdot 0) = (2, 0, -2)$; then , , , total 8.52, so P(Paris) = 7.39 / 8.52 = 0.87.
Level 3: in Python
import math
# rows: Paris, London, Rome
W_U = [[1, 0], [0, 1], [-1, 0]]
h = [0.2, 0.3]
writes = [[0.6, 0.0], [1.2, -0.3]]
# a residual layer adds its write to the stream
for w in writes:
h = [a + b for a, b in zip(h, w)]
[round(x, 2) for x in h] # → [2.0, 0.0]
# W_U h: one score per word
scores = [sum(u * x for u, x in zip(row, h)) for row in W_U]
[round(s, 2) for s in scores] # → [2.0, 0.0, -2.0]
# softmax
exps = [math.exp(s) for s in scores]
[round(e / sum(exps), 2) for e in exps] # → [0.87, 0.12, 0.02]
Figure 6 · Diagram
flowchart LR
E["Embedding"] --> H0(("h₀")) --> L1["Layer 1<br/>adds its write"] --> H1(("h₁")) --> L2["Layer 2<br/>adds its write"] --> H2(("h₂")) --> OUT["Final norm + W_U<br/>(the real output)"]
H0 -.-> LENS0["lens: norm + W_U<br/>London 0.40"]
H1 -.-> LENS1["lens: norm + W_U<br/>Paris 0.55"]
H2 -.-> LENS2["lens: norm + W_U<br/>Paris 0.87"]
primer.ml.transformer.TinyGPT, they do).On a model where we know the answer. LandmarkModel is a two-layer,
one-head transformer built by hand to complete "Eiffel is in" with
"Paris". Layer 1 is an MLP that looks up a fact at every word (landmark
in, city out); layer 2 is an attention head that lets the last word, "in",
find the landmark and copy its city.
Figure 7 · Diagram
flowchart LR T["Eiffel · is · in"] --> EMB["Embed each word"] EMB --> MLP["Layer 1: MLP at every word<br/>Eiffel → writes 'Paris'"] MLP --> ATT["Layer 2: attention head<br/>'in' looks for the landmark,<br/>copies its city"] ATT --> U["Output at 'in':<br/>Paris 3, Rome 0, London 0"]
Figure 5 · Drawn from the lesson's code
Left, the worked example's probabilities for Paris, London and Rome after the embedding, layer 1 and layer 2. Right, a 3 by 3 grid of the lens's probability of Paris by layer and word: the Eiffel column turns dark after layer 1, the in column only after layer 2
The logit lens was first described for GPT-2 in 2020, where middle layers already "guess" the next word surprisingly well. In some models the early layers read as nonsense through the final output layer, because they do not yet speak its "language"; the tuned lens fixes this by training a small translator for each layer before applying the output layer.
In code: logit_lens applies an output layer to any residual states, worked_lens runs the table above, residual_stream and tinygpt_logit_lens do the same for primer.ml.transformer.TinyGPT, and lens_map builds the grid for LandmarkModel.
Chapter 4
Activation patching: which activations cause the answer?
Everyday picture Two cars of the same model sit side by side: one starts, one doesn't. You move parts from the good car into the bad one, one part at a time, and try the ignition after each swap. The part whose swap makes the bad car start is the part that mattered. Activation patching (also called causal tracing) does this with a model's internal activations.
A tiny worked example. Run the landmark model twice:
- the clean prompt "Eiffel is in": scores Paris 3, Rome 0, so the logit difference Paris − Rome is +3;
- the corrupted prompt "Colosseum is in": Paris 0, Rome 3, so the difference is −3.
Now run the corrupted prompt again, but at one chosen place overwrite the activation with the one from the clean run, and see how much of the gap between −3 and +3 comes back:
- patch the MLP output at the landmark's position: the difference jumps back to +3, so 100% of the answer is restored;
- patch the state at "is": it stays at −3, 0% restored ("is" is the same word in both prompts, so its state carries nothing about the landmark).
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| logit difference: the correct answer's score minus the wrong answer's score (Paris − Rome) | ||
| on the clean prompt, nothing patched | +3 | |
| on the corrupted prompt, nothing patched | −3 | |
| on the corrupted prompt, with one clean activation copied in | +3, −3, or anything between | |
| restored | the fraction of the gap that one patch closes: 0 = nothing, 1 = everything | 1, 0 |
In words: "how far the patch moved the answer, as a share of the full distance from the corrupted answer to the clean one."
With the numbers: patching the landmark's MLP output gives (3 − (−3)) / (3 − (−3)) = 6 / 6 = 1. Patching "is" gives (−3 − (−3)) / 6 = 0. A patch that left the difference at 0 would give (0 − (−3)) / 6 = 0.5: half the answer restored.
Level 3: in Python
clean, corrupt = 3.0, -3.0
# restored = (LD_patched − LD_corrupt) / (LD_clean − LD_corrupt)
def restored(patched):
return (patched - corrupt) / (clean - corrupt)
# the landmark's MLP output, patched in
restored(3.0) # → 1.0
# the state at "is", patched in
restored(-3.0) # → 0.0
# a patch that only reaches a tie
restored(0.0) # → 0.5
Why a difference of two scores, rather than the probability of Paris? Because the corrupted prompt differs from the clean one in exactly one fact, the difference measures exactly that fact, and it moves smoothly, where a probability can saturate near 0 or 1 and hide a change.
Figure 9 · Diagram
flowchart TB
subgraph C["1. Clean run: 'Eiffel is in'"]
c1["save every activation"] --> c2["Paris − Rome = +3"]
end
subgraph K["2. Corrupted run: 'Colosseum is in'"]
k1["Paris − Rome = −3"]
end
subgraph P["3. Patched run: corrupted prompt,<br/>ONE activation taken from the clean run"]
p1["Paris − Rome = ?"]
end
c1 -- "copy one activation" --> P
P --> R["fraction restored =<br/>(? − (−3)) / (3 − (−3))"]
K --> R
C --> R
Figure 8 · Drawn from the lesson's code
A 3 by 3 grid, layers by positions, of the fraction of the answer restored by one patch: 1 at the Eiffel column for the embedding and layer 1, 1 at the in column after layer 2, 0 everywhere else
Compare with the lens grid above. After layer 2, the lens reads "Paris" at the Eiffel position with probability 0.995, yet patching that state restores 0.00: after the last layer nothing reads that position any more. The information is there, and it no longer matters. Readable is not the same as used, again.
Why it matters in practice. This is how researchers located where GPT-style models store facts: causal tracing on prompts like ours found factual recall concentrated in middle-layer MLPs at the subject's last token, and then used that to edit single facts. The same method, one site at a time, has traced whole circuits, such as the one GPT-2 small uses to fill in "When Mary and John went to the store, John gave a drink to" → "Mary".
In code: LandmarkModel is the hand-built model and LandmarkModel.run accepts patches; fraction_restored is the formula and patching_map patches every layer and position in turn.
Chapter 5
Superposition: more features than neurons
Everyday picture Back to the band, but now five instruments play through two speakers. The engineer pans each instrument to its own position around the room: hard left, front right, back left, and so on. When one instrument plays alone, you can tell which one by where the sound comes from. When two play at once, the positions blur together and you might mistake the pair for a third instrument. The trick works because in this band, most of the time, only one instrument is playing.
Models face the same squeeze. There are far more concepts in the world than neurons in a layer, but in any one sentence almost all of them are absent: features are sparse. So models store more features than they have dimensions, as directions that are nearly perpendicular rather than exactly. That is superposition.
A tiny worked example. Put five features in two neurons, spread 72° apart like the points of a pentagon. Feature 's direction is (cos 72k°, sin 72k°): feature 0 is (1, 0), feature 1 is (0.309, 0.951), and so on. Switch on feature 0 alone, at strength 1. The hidden state is (1, 0). Reading every feature back with a dot product gives
| Feature | Angle from feature 0 | Read-back (cos of the angle) | After adding −0.31 and ReLU |
|---|---|---|---|
| 0 | 0° | 1.000 | 0.69 |
| 1 | 72° | 0.309 | 0 |
| 2 | 144° | −0.809 | 0 |
| 3 | 216° | −0.809 | 0 |
| 4 | 288° | 0.309 | 0 |
Five directions can't all be perpendicular in two dimensions, so reading feature 0 leaks 0.309 into each neighbour: interference. The fix is a small negative bias and a ReLU (which turns negatives into zero): the leak of 0.309 falls below the bias of 0.31 and is filtered to 0, and the real feature survives at 0.69. The cost comes when two neighbours are on at once: each then reads 1 + 0.309 − 0.31 = 0.999 instead of 0.69, too high. Sparsity is a bet that such collisions are rare.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the true features: one strength per feature | (1, 0, 0, 0, 0) | |
| the squeeze: 2 rows (neurons) by 5 columns (features); column is feature 's direction | the pentagon | |
| the hidden state: 2 numbers holding 5 features | (1, 0) | |
| transposed (rows and columns swapped): reads each feature back with a dot product | ||
| one bias per feature, learned; negative values filter small leaks | −0.31 each | |
| ReLU | keep positives, turn negatives into 0 | |
| the features read back out | (0.69, 0, 0, 0, 0) | |
| feature 's direction dotted with itself (its squared length) | 1 | |
| how much feature leaks into feature : 0 if perpendicular | 0.309 for neighbours | |
| add up over every other feature |
In words: "squeeze the features into a few neurons, read each one back by dotting with its direction, subtract a threshold and drop anything negative. What you read for feature is its own strength, plus a leak from every other active feature, minus the threshold."
With the numbers: with feature 0 alone on, $\hat{x}_1 = \text{ReLU}(1 \cdot 0 + 0.309 \cdot 1 - 0.31) = \text{ReLU}(-0.001) = 0$ and . With features 0 and 1 both on, .
Level 3: in Python
import math
# feature k's direction: (cos 72k°, sin 72k°)
W = [[math.cos(math.radians(72 * k)), math.sin(math.radians(72 * k))] for k in range(5)]
def read_back(x, bias, relu=True):
# h = W x: two numbers
h = [sum(W[k][n] * x[k] for k in range(5)) for n in range(2)]
# Wᵀ h + b: one number per feature
z = [W[k][0] * h[0] + W[k][1] * h[1] + bias for k in range(5)]
# ReLU keeps positives and turns negatives into 0
return [round(max(0.0, v) if relu else v, 3) for v in z]
# the leak: feature 0 alone, no bias, no ReLU
read_back([1, 0, 0, 0, 0], bias=0, relu=False) # → [1.0, 0.309, -0.809, -0.809, 0.309]
# the bias and ReLU filter it
read_back([1, 0, 0, 0, 0], bias=-0.31) # → [0.69, 0.0, 0.0, 0.0, 0.0]
# two neighbours on: each reads too high
read_back([1, 1, 0, 0, 0], bias=-0.31) # → [0.999, 0.999, 0.0, 0.0, 0.0]
Figure 11 · Diagram
flowchart LR X["5 features<br/>x₀ … x₄<br/>(mostly zero)"] --> W["squeeze: W<br/>(2 × 5)"] W --> H["hidden state<br/>2 neurons"] H --> WT["read back: Wᵀ<br/>(5 × 2)"] WT --> B["+ bias b<br/>(negative)"] B --> R["ReLU"] R --> XH["5 features<br/>read back"]
Training it. Let the model learn and itself, by gradient descent on the reconstruction error, with features that matter less and less (feature is weighted ). Compare two worlds:
- Dense: every feature is on in every example. The model keeps the two most important features, perpendicular, at length 1.00, and gives the other three length 0: it simply drops them.
- Sparse: each feature is on only 5% of the time. The model keeps all five, 72° apart, at length about 1.1, and learns a negative bias of about −0.23 to filter the leaks.
Figure 10 · Drawn from the lesson's code
Three panels. Left, dense features: two perpendicular arrows and three near-zero ones. Middle, sparse features: five arrows spread evenly around a pentagon. Right, bars of how much each of five features moves neuron 1 and neuron 2
That last point is why looking at single neurons so often fails. A neuron that responds to several unrelated features is polysemantic. Vision researchers found neurons that respond to cat faces and to the fronts of cars; language models are full of neurons like that. Superposition explains why: when features outnumber neurons, the features cannot line up with the neurons, so every neuron is a mixture.
In code: pentagon builds the five directions, superposition_readout is the formula, train_superposition learns W and b with gradients from superposition_loss_and_grads (checked against finite_difference_gradient), and features_a_neuron_responds_to lists a neuron's features.
Chapter 6
Sparse autoencoders: getting the features back
Everyday picture A sound engineer receives the two-speaker recording of the five-instrument band, with no notes on who played when. She knows one thing about this band: usually only one instrument plays at a time. Many different scores could produce the same sound, but she writes down the one that uses the fewest instruments. That preference is what lets her recover the real parts rather than some arbitrary mixture.
A sparse autoencoder (SAE) does this for a model's hidden states. It learns a dictionary: many more directions than the layer has neurons (it is overcomplete), trained so that every hidden state can be rebuilt from just a few of them. Each learned direction is a candidate feature, and each one's strength on an input is its latent activation.
A tiny worked example. The hidden state is h = (0.8, 0.6) and the dictionary has three directions, (1, 0), (0, 1) and (0.8, 0.6). Two codes rebuild h perfectly:
| Code (strength of each direction) | Rebuilt | Error | Penalty with λ = 0.1 | Loss |
|---|---|---|---|---|
| A: (0.8, 0.6, 0), two directions | (0.8, 0.6) | 0 | 0.1 × (0.8 + 0.6) = 0.14 | 0.14 |
| B: (0, 0, 1), one direction | (0.8, 0.6) | 0 | 0.1 × 1.0 = 0.10 | 0.10 |
Rebuilding alone can't choose between them. The penalty on the total
strength, the L1 penalty (primer.ml.regularization), prefers the code
that uses one direction. One side effect: the best strength for direction
3 is not quite 1. With strength , the loss is , which
is lowest at (loss 0.0975): the penalty always pulls strengths
a little toward zero, known as shrinkage.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| a hidden state from the model under study | (0.8, 0.6) | |
| how many latents (dictionary directions) the SAE has, usually far more than neurons | 3 | |
| , | the encoder's weights and biases: turn a hidden state into latent strengths | |
| the code: one strength per latent, ReLU keeps them ≥ 0 and mostly exactly 0 | (0, 0, 1) | |
| the decoder: column is latent 's direction, kept at length 1 | (1, 0), (0, 1), (0.8, 0.6) | |
| the decoder bias: the typical hidden state, subtracted before encoding and added back after | (0, 0) | |
| the rebuilt hidden state | (0.8, 0.6) | |
| squared length of the rebuild error: add up the squares of its entries | 0 | |
| the sparsity penalty's strength, chosen by you | 0.1 | |
| the size of latent 's strength | ||
| the loss to minimize: rebuild error plus penalty | 0.10 |
In words: "encode the hidden state into many non-negative strengths, decode it back as a weighted sum of dictionary directions, and pay for both the rebuild error and the total strength used."
With the numbers: for code B, $\hat{h} = 1 \cdot (0.8, 0.6) = (0.8, 0.6)0.1 \times 1 = 0.10L = 0.10$. For code A the penalty is .
Level 3: in Python
dictionary = [[1, 0], [0, 1], [0.8, 0.6]]
h = [0.8, 0.6]
lam = 0.1
def loss(code):
# ĥ = Σ f_i · direction_i
rebuilt = [sum(f * d[n] for f, d in zip(code, dictionary)) for n in range(2)]
# ‖h − ĥ‖² + λ Σ |f_i|
return round(sum((a - b) ** 2 for a, b in zip(h, rebuilt)) + lam * sum(abs(f) for f in code), 4)
loss([0.8, 0.6, 0]) # → 0.14
loss([0, 0, 1.0]) # → 0.1
# shrinkage: a slightly weaker code is cheaper still
loss([0, 0, 0.95]) # → 0.0975
Figure 13 · Diagram
flowchart LR H["hidden state h<br/>(d numbers)"] --> ENC["encoder<br/>ReLU(W_e(h − b_d) + b_e)"] ENC --> F["code f<br/>(m ≫ d numbers,<br/>almost all exactly 0)"] F --> DEC["decoder<br/>W_d f + b_d"] DEC --> HH["rebuilt ĥ"] HH --> LOSS["loss = ‖h − ĥ‖² + λ Σ|f|"] H --> LOSS F --> LOSS
On the superposition toy. Take 20,000 hidden states from the pentagon model (each feature on 5% of the time), train an SAE with 5 latents and λ = 0.3, and compare the learned decoder columns with the five planted directions. Every planted feature is matched by a latent with cosine similarity above 0.99 (1 would be the same direction), and an example with a feature on lights up about 1.07 latents on average. The price is shrinkage: 11% of the variance goes unexplained. With a tiny penalty, λ = 0.01, the SAE rebuilds the data perfectly (0.0% unexplained), yet its worst match to a real feature is only 0.80: any two directions can rebuild a two-dimensional space, so without sparsity pressure nothing forces the dictionary to find the model's actual features.
Figure 12 · Drawn from the lesson's code
Left, grey dots of hidden states lying mostly along five rays, with thick grey planted directions and red learned latent arrows lying on top of them. Right, as the penalty grows from 0.003 to 1, the worst feature match peaks near 0.99 at a penalty of 0.3, while the unexplained variance climbs from 0 to above 0.5
Why it matters in practice. This is the tool that turned interpretability from "a few hand-picked neurons" into something that scales. Anthropic's Towards Monosemanticity (2023) trained SAEs on a small transformer and found thousands of features that each fire for one recognizable thing. Scaling Monosemanticity (2024) did the same inside a production model, Claude 3 Sonnet, and found features for concepts such as the Golden Gate Bridge. Turning that one feature up made the model bring up the bridge in almost every answer, which is the causal check that the feature means what it seems to mean.
In code: sae_objective computes the loss for one proposed code, SparseAutoencoder holds the encoder and decoder with its hand-derived SparseAutoencoder.loss_and_grads, train_sae fits it, and match_features compares learned directions with planted ones.
Chapter 7
What these tools cannot (yet) show
Everyday picture A brain scan shows which regions light up while a person reads, but not the sentence they are thinking. Every tool in this lesson is like that: a real instrument that measures something true and partial.
A tiny worked example. Each of our own toys already showed a limit:
| What we saw | The limit it shows |
|---|---|
| A probe read tense at 98%, and the output ignored tense completely | Probes find information, not use |
| The lens read "Paris" at a position where patching restored 0% | Reading an activation doesn't mean anything downstream reads it |
| The weak SAE rebuilt everything perfectly and still found the wrong directions | A good rebuild score doesn't mean the features are real |
| The strong SAE left 11% of the variance unexplained | Some of the model's activity is in no feature the dictionary found |
| We knew the landmark model's circuit because we built it | In a real model, nobody has the answer key |
Figure 14 · Diagram
flowchart LR A["Correlation<br/>probes, the logit lens:<br/>'this information is here'"] --> B["Intervention<br/>patching, steering:<br/>'changing this changes the output'"] B --> C["Mechanism<br/>a circuit explained end to end:<br/>'this is how the output is computed'"]
The main open problems, in plain words:
- Choices shape answers. Patching results depend on how the corrupted prompt is built (swap one word? add noise?). SAE features depend on the dictionary's size and λ: a bigger dictionary can split one feature into several finer ones.
- Redundancy hides importance. Patching one site at a time can miss parts that back each other up. Models have been observed to partly repair themselves when one component is knocked out, so "patching this restores 0%" does not always mean "this plays no role".
- Coverage. The unexplained part of an SAE's rebuild is not noise to ignore: it is model behaviour no feature describes yet.
- Scale and labour. A circuit that explains one behaviour of a small model can take researchers weeks. No one has a complete account of how a frontier model produces any long answer.
- Labels are human guesses. Naming a feature "Golden Gate Bridge" summarizes the inputs that make it fire; checking that the name is right needs interventions like the one above.
Why it matters in practice. Treat these tools as evidence, stacked
next to behavioural evaluations (primer.agents.evals), never as a
certificate. A probe or a lens is a cheap first look; a patching or
steering experiment is the minimum for a causal claim; and a clean result
on a toy, like every result in this lesson, is where understanding starts,
not where it ends.
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1What does it mean to say a feature is a "direction" rather than a neuron?Think it through, then reveal
The feature's presence is stored as a pattern across many neurons: the hidden state moves along a particular direction in proportion to how strongly the feature is present. You read it with a dot product against that direction. Any single neuron is typically a mix of several features.
Question 2A linear probe reads a property from layer 12 at 95% accuracy. What can you conclude, and what can't you?Think it through, then reveal
You can conclude the property is linearly decodable from layer 12, as long as the 95% was on held-out examples and clearly beats a control task with random labels. You can't conclude the model uses it. For that you need an intervention: remove or change that information and see whether the output changes.
Question 3How does the logit lens work, and why does it agree with the model at the last layer?Think it through, then reveal
It takes the residual stream after some layer and applies the model's own final normalization and unembedding, turning it into next-token probabilities. At the last layer that is exactly the computation the model itself performs, so the two must match. Earlier layers are read "as if the model stopped there".
Question 4Describe an activation patching experiment, and what the "fraction restored" means.Think it through, then reveal
Run a clean prompt and save its activations; run a corrupted prompt that changes one fact; then rerun the corrupted prompt with one activation replaced by its clean value. The fraction restored is how much of the clean-minus-corrupted logit difference the patch brings back: 1 means that activation alone carries the fact, 0 means it carries none of it.
Question 5Why do models use superposition, and why does it make neurons hard to interpret?Think it through, then reveal
There are more useful features than neurons. When features are rarely active at the same time, the model can store them as nearly perpendicular directions and filter the small interference with a bias and ReLU. The directions can't line up with the neurons, so each neuron responds to several unrelated features: it is polysemantic.
Question 6Why does a sparse autoencoder need the L1 penalty? What goes wrong if λ is too small or too large?Think it through, then reveal
Many dictionaries rebuild the data equally well; the penalty picks the one where each input uses few latents, which pushes latents onto the real features. Too small and the SAE rebuilds perfectly with meaningless directions; too large and it shrinks activations and leaves much of the model's activity unexplained.
Primary sources
The papers behind this lesson
Introduced linear probes as a way to measure what each layer of a network makes linearly available.
The paper ↗Showed that probes can succeed by memorizing, and introduced control tasks and selectivity to tell the two apart.
The paper ↗Formalized the logit lens and fixed its failures on early layers by training a small translator for each layer.
The paper ↗Introduced causal tracing, found factual recall in middle-layer MLPs at the subject's last token, and edited single facts there.
Read the annotated companion →The paper ↗Used patching to reverse-engineer a complete circuit for one behaviour of a real language model.
The paper ↗Showed with small ReLU models that sparse features are stored in superposition, and when and how the geometry changes.
Read the annotated companion →The paper ↗Trained sparse autoencoders on a small transformer and found thousands of interpretable features hidden in polysemantic neurons.
Read the annotated companion →The paper ↗Scaled sparse autoencoders to a production model and steered its behaviour through the features they found.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Olah et al., Zoom In: An Introduction to Circuits (Distill, 2020): https://distill.pub/2020/circuits/zoom-in/
- Elhage et al., A Mathematical Framework for Transformer Circuits (2021): https://transformer-circuits.pub/2021/framework/index.html
- Elhage et al., Toy Models of Superposition (2022): https://transformer-circuits.pub/2022/toy_model/index.html
- Bricken et al., Towards Monosemanticity (2023): https://transformer-circuits.pub/2023/monosemantic-features/index.html
- Templeton et al., Scaling Monosemanticity (2024): https://transformer-circuits.pub/2024/scaling-monosemanticity/index.html
- Cunningham et al., Sparse Autoencoders Find Highly Interpretable Features in Language Models (2023): https://arxiv.org/abs/2309.08600
- Belinkov, Probing Classifiers: Promises, Shortcomings, and Advances (2021): https://arxiv.org/abs/2102.12452
- Meng et al., Locating and Editing Factual Associations in GPT (2022): https://arxiv.org/abs/2202.05262
- Belrose et al., Eliciting Latent Predictions from Transformers with the Tuned Lens (2023): https://arxiv.org/abs/2303.08112
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.