At a glance
Key takeaways
- Chain of thought: every written token is another forward pass, so writing steps gives a fixed-depth model more serial computation, and the text is its working memory.
- Test-time compute: spend more tokens at answer time, either one longer chain or many chains; accuracy often grows with the logarithm of the budget until the problems run out.
- Self-consistency: sample several chains and vote on the final answer; it helps when the right answer is the most common one and mistakes are independent, and stalls when they are shared.
- Verifiers: outcome checks score the answer, process checks score each step; best-of-n with a good checker approaches pass@n = 1 − (1 − p)ⁿ.
- RL with verifiable rewards: reward right final answers, compare each attempt with its group (GRPO), and longer, self-checking reasoning emerges because it pays; a length penalty keeps it from overthinking.
- Cost and limits: thinking is paid per token, slips compound over long chains, and the written chain is not a guaranteed account of why the model answered.
Level 2
How it works, from scratch
Ask someone to multiply 37 by 48 in their head, instantly, and they will probably guess. Hand them a pencil and a minute and they will get it right. They didn't get smarter in that minute. They got room: somewhere to write 37 × 8 = 296 and 37 × 40 = 1,480, so each small step could lean on the last.
A language model is in the same position. It can do a fixed amount of work for each token it writes: one pass through all of its layers. A reasoning model is a model that has learned to use the pencil. Before it answers, it writes out intermediate steps, checks them, backs up when one is wrong, and only then commits to an answer. This lesson builds each piece from scratch:
- why writing steps helps at all (chain of thought);
- why more thinking can buy more accuracy (test-time compute);
- sampling several attempts and voting on the answer (self-consistency);
- checking attempts, either the final answer or every step (verifiers);
- how reinforcement learning teaches a model to think this way;
- what thinking costs, and where it still fails.
Chapter 1
Why thinking out loud helps
Everyday picture Picture a clerk who, each time they glance at a page, can do exactly one addition and write down one number. Hand them a column of four numbers and allow one glance, and they have to guess. Allow three glances and a margin to write in, and they get it right every time. The margin is the whole trick: what they write on one glance is there to read on the next.
Tiny example Our toy model is that clerk. Every token it writes costs one forward pass (one run of the whole model to produce one token), and each forward pass can do one addition. Ask it for 3 + 5 + 8 + 2, which needs three additions.
| Pass | Answer at once (1 token) | Think first, then answer |
|---|---|---|
| 1 | adds 3 + 5 = 8, has no pass left for 8 and 2, estimates them at 4.5 each: 8 + 9 = 17 ✗ | writes 3+5=8 |
| 2 | reads 8, writes 8+8=16 |
|
| 3 | reads 16, answers 16 + 2 = 18 ✓ |
The weights are identical in both columns. The only difference is that the second column wrote its running totals down, where the next pass could read them. Writing intermediate steps before the answer is called chain of thought.
Figure 1 · Diagram
flowchart LR
subgraph D["Answer at once: 1 token"]
direction LR
q1["3+5+8+2 ="] --> p1["pass 1<br/>3+5 = 8<br/>no pass left"] --> a1["17, a guess"]
end
subgraph C["Think first: 3 tokens"]
direction LR
q2["3+5+8+2 ="] --> s1["pass 1<br/>writes 3+5=8"] --> s2["pass 2<br/>reads 8<br/>writes 8+8=16"] --> s3["pass 3<br/>reads 16<br/>answers 18"]
end
The rule behind the table: the number of steps a model can do one after another grows with the number of tokens it writes.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| serial steps: how many steps the model can chain, each using the result of the one before | 1 or 3 | |
| serial steps one forward pass can do; for a transformer, roughly its number of layers | 1 addition in the toy | |
| tokens generated, the answer included; each token is one forward pass | 1 at once, 3 thinking first | |
| multiply |
In words: "the length of the chain of steps a model can work through equals the steps per token times the number of tokens it writes."
With the numbers: the toy has and the sum needs 3 additions. Answering at once, , so : two additions short, so it guesses. Thinking first, , so : exactly enough. A 32-layer model answering at once has at most 32 layers of one-after-another work; with a 1,000-token chain of thought it has up to 32,000.
Level 3: in Python
# the toy: 1 addition per pass; 3 + 5 + 8 + 2 needs 3 additions
L = 1
# answering at once: T = 1, so S = 1, not enough
L * 1 # → 1
# the one pass adds 3 + 5, then estimates the unreached 8 and 2 at 4.5 each
(3 + 5) + round(4.5 * 2) # → 17
# thinking first: T = 3, so S = 3, exactly enough
L * 3 # → 3
t1 = 3 + 5 # → 8
t2 = t1 + 8 # → 16
t2 + 2 # → 18
# a 32-layer model: answering at once, then after a 1,000-token chain of thought
32 * 1 # → 32
32 * 1000 # → 32000
A layer is not literally one addition, so is a rough count, but the shape of the rule holds: a model with a fixed depth can only do a fixed amount of one-after-another work per token, and problems that need more must spread it over more tokens. Li et al. (2024) proved a version of this for transformers: with enough chain-of-thought tokens they can solve inherently serial problems that a fixed-depth transformer answering at once cannot.
Figure 2 · Drawn from the lesson's code
Answering in one token, the toy is right only on one-addition sums and about 5% of the rest; writing steps, it is always right, at one token per addition
That price is real: each token costs roughly floating-point operations
for a model with parameters (see primer.ml.inference), so a
1,000-token chain costs a thousand answers' worth of compute. In practice
this is why simply asking a model to "think step by step" (Kojima et al.,
2022) or showing it worked examples with steps (Wei et al., 2022) improved
accuracy on arithmetic and logic puzzles, and why every reasoning model
writes a long chain before its answer.
In code: solve is the toy model: it spends one forward pass per written step, one on the answer, and estimates anything it never reached. It returns an Attempt holding the steps, the answer and the tokens used.
Chapter 2
Test-time compute: buying accuracy with tokens
Everyday picture An exam has questions of very different difficulty. Give yourself one minute per question and you finish only the easiest. Two minutes, and the next tier falls. Each doubling of time unlocks one more tier, until you can finish everything and extra time buys nothing.
Tiny example Test-time compute is computation spent while answering, as opposed to while training. The simplest dial is a thinking budget: the most tokens the model may write before it must answer. Take six problems that need 1, 2, 4, 8, 16 and 32 additions, and a budget of 8 tokens. Four of them fit (1, 2, 4, 8). The other two must be guessed, and say a guess is right 10% of the time. Accuracy is 4/6 + 2/6 × 0.1 = 0.70.
There are two ways to spend test-time compute, and the rest of this lesson uses both.
Figure 3 · Diagram
flowchart LR P[Problem] --> LONG["Think longer<br/>one chain, bigger budget"] P --> WIDE["Think wider<br/>n chains side by side"] LONG --> A1[Answer] WIDE --> PICK["Pick one:<br/>vote or verifier"] --> A2[Answer]
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the thinking budget, in tokens | 8 | |
| the share of problems that need at most steps, so fit in the budget | 4/6 = 0.667 | |
| the share that doesn't fit and must be guessed | 2/6 = 0.333 | |
| the chance a guess happens to be right | 0.1 | |
| the share of problems answered correctly with budget | 0.70 |
In words: "accuracy is the share of problems that fit in the budget, plus a lucky share of the ones that don't."
With the numbers: , so . Doubling to lets a fifth level fit: , . Every doubling adds the same , until fits everything and .
Level 3: in Python
needs = [1, 2, 4, 8, 16, 32]
g = 0.1
B = 8
# F(B): the share of problems that fit
F = sum(d <= B for d in needs) / len(needs)
round(F, 3) # → 0.667
# acc(B) = F(B) + (1 − F(B)) · g
round(F + (1 - F) * g, 3) # → 0.7
# each doubling of B lets one more of the six levels fit
round((1 / len(needs)) * (1 - g), 3) # → 0.15
Figure 4 · Drawn from the lesson's code
Accuracy climbs 0.15 per doubling of the budget, from 0.25 at 1 token to 1.0 at 32, and the simulated sums land on the formula; tokens actually spent level off at about 10
The straight line on a log axis is built into this toy, because its difficulties are spaced by doubling. Real problem sets are also spread over many scales of difficulty, and published reasoning models show the same shape over a useful range: accuracy rising roughly in proportion to the logarithm of thinking tokens, then leveling off. Snell et al. (2024) found that spending test-time compute adaptively (more on harder prompts) beats spending it evenly, and that on problems a small model can sometimes solve, extra test-time compute can stand in for a much larger model. In practice, model APIs expose this dial as a "reasoning effort" setting or a maximum number of thinking tokens.
In code: budget_accuracy is the formula, and budget_sweep runs solve on random sums at each budget and reports accuracy and tokens actually spent.
Chapter 3
Sampling many and voting: self-consistency
Everyday picture Unsure of an answer, you ask five friends separately. If four of them say the same thing, you trust it. Each friend can be wrong, but it is unlikely that most of them are wrong in the same way, unless they all read the same wrong article.
Tiny example A model that samples its tokens (see primer.ml.inference)
writes a different chain each time. Ask the toy's noisy cousin for
3 + 5 + 8 + 2 three times and it answers 18, 17, 18. Keep only the final
answers and take the most common one: 18. Sampling several chains of thought
and returning the most common final answer is self-consistency (Wang et
al., 2022), also called majority voting.
Figure 5 · Diagram
flowchart LR Q["3+5+8+2 = ?"] --> S1["chain 1 ends in 18"] & S2["chain 2 ends in 17"] & S3["chain 3 ends in 18"] S1 & S2 & S3 --> V["count final answers<br/>18: two votes, 17: one"] --> A["answer 18"]
How much does voting help? Take the simplest case, a yes/no question: every wrong vote lands on the same wrong answer, and each vote is right with the same chance , independently of the others.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| number of sampled votes (odd, so there are no ties) | 3 | |
| chance one vote is right | 0.6 | |
| how many of the votes are right | 2 or 3 | |
| rounded up: the smallest number of votes that wins | 2 | |
| " choose ": how many ways to pick which of the votes are the right ones | ||
| the chance of one particular pattern: those right, the other wrong | ||
| add up over every winning count of right votes | ||
| the chance the majority is right | 0.648 |
In words: "the vote is right when at least half the votes are right, so add up the chance of every such count: the number of ways to get that many right, times the chance of each way."
With the numbers: , . Two right: $3 \times 0.6^2 \times 0.4 = 0.4320.6^3 = 0.216$. Together, 0.648, up from 0.6 for one vote. Five votes give 0.683. But a solver right only 40% of the time gets worse with five votes: 0.317. Voting amplifies whichever answer is most common, right or wrong. (This is the Condorcet jury theorem, from 1785.)
Level 3: in Python
import math
p, n = 0.6, 3
# one term per winning count k = 2, 3
terms = [math.comb(n, k) * p**k * (1 - p)**(n - k) for k in range(math.ceil(n / 2), n + 1)]
[round(t, 3) for t in terms] # → [0.432, 0.216]
round(sum(terms), 3) # → 0.648
# five votes
round(sum(math.comb(5, k) * 0.6**k * 0.4**(5 - k) for k in range(3, 6)), 3) # → 0.683
# a 40% solver on a yes/no question: voting makes it worse
round(sum(math.comb(5, k) * 0.4**k * 0.6**(5 - k) for k in range(3, 6)), 3) # → 0.317
Figure 6 · Drawn from the lesson's code
Votes of an 80% or 60% solver climb towards 1, a 40% solver on a yes/no question sinks towards 0, but the same 40% solver climbs past 0.9 when its wrong answers scatter over ten values
When voting fails: shared mistakes. Everything above assumed the votes were independent. Samples from one model are not: they share its training, its blind spots and its reading of the question. Model that as a shared draw: each sample copies one common answer with probability (rho), and otherwise answers on its own. Each single sample is still right 60% of the time, so only the correlation changes.
Figure 7 · Drawn from the lesson's code
With independent samples 31 votes reach nearly 1.0; with rho = 0.3 the vote levels off near 0.86; with rho = 0.6 or 0.9 it stays at about 0.6, no better than one sample
In practice, self-consistency gives a solid gain for a few samples and then flattens, because a model's mistakes on a question are correlated. Diversity helps (different prompts, different models, a tool that computes instead of guessing), and voting only works where answers can be compared exactly.
In code: majority_vote picks the most common answer, majority_accuracy is the formula (with ties on even counted as a coin flip), and correlated_vote_accuracy simulates votes with scattered wrong answers and a shared-mistake rate.
Chapter 4
Verifiers: checking the answer, or checking every step
Everyday picture One teacher looks only at the boxed answer at the bottom of the page. Another marks every line of working. The first can't tell a lucky guess from understanding, and when the answer is wrong, can't tell you where you went wrong. The second can do both.
Tiny example A verifier is anything that scores a candidate solution. Here are two chains the toy wrote for 3 + 5 + 8 + 2 (true answer 18), each with slips:
| Chain | Final answer | Outcome check: is it 18? | Step check: first line that is false |
|---|---|---|---|
3+5=8, 8+8=17, 17+2=19 |
19 | reward 0 | step 2: 8 + 8 is 16 |
3+5=9, 9+8=16, 16+2=18 |
18 | reward 1 | step 1: 3 + 5 is 8 |
An outcome verifier looks only at the final answer: either a learned outcome reward model (ORM) that predicts whether it is right, or a check against a known answer. It gives the second chain full marks, although two slips just happened to cancel. A process verifier, or process reward model (PRM), scores every step. It catches the second chain's slip and points at exactly where the first one went wrong.
Figure 8 · Diagram
flowchart LR
subgraph O["Outcome verifier"]
direction LR
o1["3+5=8"] --> o2["8+8=17"] --> o3["17+2=19"] --> oc{"final 19<br/>equals 18?"} --> orr["reward 0<br/>but where did it go wrong?"]
end
subgraph P["Process verifier"]
direction LR
p1["3+5=8<br/>ok"] --> p2["8+8=17<br/>wrong"] --> p3["17+2=19<br/>ok"] --> pr["first bad step: 2"]
end
Verifiers power best-of-n: sample chains, keep the one the verifier scores highest. With a perfect verifier, best-of-n is right whenever any of the samples is right, a number called pass@n.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| chance one sample is right | 0.3 | |
| number of samples drawn | 5 | |
| chance one sample is wrong | 0.7 | |
| chance all are wrong, multiplying because samples are independent | ||
| chance at least one of the is right | 0.832 |
In words: "the chance that at least one sample is right is one minus the chance that every sample is wrong."
With the numbers: a model right 30% of the time is wrong on five tries in a row with chance , so pass@5 = 0.832. A 30% model with a perfect checker and five tries beats an 80% model on one try. In the experiment below, each chain has 5 additions that each slip with chance 0.2, so a chain is clean with chance .
Level 3: in Python
p, n = 0.3, 5
# (1 − p)^n: every one of the five is wrong
round((1 - p) ** n, 5) # → 0.16807
round(1 - (1 - p) ** n, 5) # → 0.83193
# the experiment's chains: 5 additions, each clean with chance 0.8
round(0.8 ** 5, 3) # → 0.328
Figure 9 · Drawn from the lesson's code
With 8 chains, one sample is right 38% of the time, a majority vote 67%, best-of-8 with a step checker 96%, against a pass@8 ceiling of 99%
In practice, the best verifiers are programs: unit tests for code, an exact match for a math answer, a proof checker. Where no program can check, a learned verifier stands in. Lightman et al. (2023) trained a process reward model on 800,000 human labels of individual steps and found it picked correct solutions far more reliably than an outcome reward model. A learned verifier can be fooled, though, and more samples mean more chances to find a wrong answer it likes: Cobbe et al. (2021) saw best-of-n accuracy start to fall again after a few hundred samples.
In code: check_step is the toy's process verifier for one line, first_bad_step runs it over a chain, and outcome_reward compares only the final answer with a reference. noisy_chain writes chains whose slips carry forward, pass_at_n is the formula, and verifier_experiment compares one sample, voting, best-of-n with the step checker, and the pass@n ceiling.
Chapter 5
Learning to reason with reinforcement learning
Everyday picture A workbook with the answers printed in the back. Nobody shows you how to solve anything. You try each problem several ways, check the back, and do more of whatever worked. Over hundreds of problems you discover good habits on your own: write things down, double-check the tricky step, don't stop too early.
Tiny example Prompting a model to "think step by step" gets it to
write steps; a reasoning model is trained to write good ones. The recipe
behind models such as DeepSeek-R1 is reinforcement learning (see
primer.ml.reinforcement) with two ingredients:
- Verifiable rewards. Train on problems whose answers a program can check: math with a known final answer, code with unit tests. The reward is 1 if the final answer is right and 0 if not. Nobody grades the chain itself.
- Group comparisons. For each problem, sample a group of attempts and judge each one against its own group. With and rewards 1, 0, 0, 1, the group averages 0.5, so the two right attempts are above average (+1) and the two wrong ones below (−1). The model is nudged towards whatever the right ones did, including how they reasoned. This is the heart of GRPO (group relative policy optimization).
Figure 10 · Diagram
flowchart LR P["problem with a checkable answer"] --> G["sample G attempts,<br/>each with its own chain of thought"] G --> R["check each final answer:<br/>reward 1 or 0"] R --> A["advantage: reward minus group mean,<br/>divided by group spread"] A --> U["make above-average attempts more likely,<br/>below-average ones less"] U -->|next batch| P
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| how many attempts are sampled for one problem | 4 | |
| which attempt, from 1 to | 1, 2, 3, 4 | |
| attempt 's reward: 1 if its final answer is right, 0 if not | 1, 0, 0, 1 | |
| the average reward in the group | 0.5 | |
the standard deviation: the typical distance of a reward from the mean (see primer.notation) |
0.5 | |
| attempt 's advantage: how much better than its group it did, in units of the group's spread | +1, −1, −1, +1 |
In words: "an attempt's advantage is how far its reward sits above its group's average, measured in units of how spread out the group's rewards are."
With the numbers: rewards 1, 0, 0, 1 have mean 0.5 and standard deviation 0.5, so the advantages are (1 − 0.5)/0.5 = +1 and (0 − 0.5)/0.5 = −1. If all four attempts are right, every reward equals the mean and the spread is 0: every advantage is 0, and the problem teaches nothing. The same holds when all four are wrong. Training therefore needs problems the model solves sometimes.
Level 3: in Python
import statistics
r = [1, 0, 0, 1]
mu = statistics.mean(r) # → 0.5
sigma = statistics.pstdev(r) # → 0.5
[(r_i - mu) / sigma for r_i in r] # → [1.0, -1.0, -1.0, 1.0]
# a group that all succeeded: spread 0, nothing to learn
statistics.pstdev([1, 1, 1, 1]) # → 0.0
The toy: learning how long to think. Our policy (the model's rule for choosing what to do) makes one choice per problem: how many tokens to think for, out of 1, 2, 4, 8, 16 or 32. It keeps a separate choice for each difficulty (problems needing 1, 2, 4 or 8 steps), and it starts out preferring short answers, like a model never rewarded for thinking: 63% of the time it answers in 1 token. A chain shorter than the steps needed must guess (right 10% of the time). Otherwise each step gets as many tries as the length allows, and a try slips 10% of the time: one pass through the steps, or two (the first attempt and a recheck that catches a slip), or more. The reward is 1 for a right answer and 0 for a wrong one. That is all.
Figure 11 · Drawn from the lesson's code
Rewarded only for correct answers, mean thinking length rises from 2 to about 7.6 tokens and accuracy from 0.41 to 0.96; with a length penalty the length and accuracy settle a little lower
Figure 12 · Drawn from the lesson's code
Trained on correctness alone, the model spends about 2.5, 4, 8 and 16 tokens on problems needing 1, 2, 4 and 8 steps: at or just past twice the steps needed
In code: group_advantages is the formula, success_probability is the toy's chance of solving a problem at a given length, and train_reasoner runs the whole loop: sample a group of lengths per difficulty, reward, compute advantages, and nudge the policy.
The price of thinking
Rewarded for correctness alone, nothing ever tells the model to stop: any extra length that helps even slightly gets reinforced. Real reasoning models show this as overthinking: hundreds of tokens spent on "what is 2 + 3?". The usual remedy is to charge for length: subtract a small penalty per token from the reward. The expected reward for a chain of length on a problem needing steps (with ) becomes:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the expected (average) reward for this length | 0.763 | |
| steps the problem needs | 8 | |
| thinking length, in tokens | 16 | |
| rounded down: how many tries each step gets | 2 | |
| chance one try at a step slips | 0.1 | |
| chance every try at one step slips, so the step fails | 0.01 | |
| chance all steps come out right | 0.923 | |
| lambda: the penalty per token of thinking | 0.01 | |
| the total charge for the chain | 0.16 |
In words: "the reward for a length is how likely it is to get every step right, given the tries it allows, minus a small charge for every token."
With the numbers: for an 8-step problem, 8 tokens give ; 16 tokens give ; 32 tokens give . Sixteen wins: one recheck is worth paying for, a third is not. For a 1-step problem, 1 token gives , 2 tokens , 4 tokens . Two wins.
Level 3: in Python
eps, lam = 0.1, 0.01
# an 8-step problem: chance of success at 8, 16 and 32 tokens
[round((1 - eps ** (L // 8)) ** 8, 3) for L in (8, 16, 32)] # → [0.43, 0.923, 0.999]
# minus λL: 16 tokens wins
[round((1 - eps ** (L // 8)) ** 8 - lam * L, 3) for L in (8, 16, 32)] # → [0.35, 0.763, 0.679]
# a 1-step problem at 1, 2 and 4 tokens: 2 tokens wins
[round((1 - eps ** (L // 1)) ** 1 - lam * L, 3) for L in (1, 2, 4)] # → [0.89, 0.97, 0.96]
Now look back at the two figures. With the penalty and no division by the spread (green), training finds exactly those best lengths, 2 tokens for the easiest problems and 16 for the hardest. With GRPO's division by the spread (red), the easiest problems shrink to 1 token, and accuracy settles at 0.92, below the 0.96 of correctness alone. Why? In a group where every attempt succeeded, rewards differ only by the penalty: 0.99 for 1 token, 0.98 for 2. Their spread is tiny, so dividing by it inflates that 0.01 difference into advantages of +1 and −1, as loud as the difference between solved and failed. The penalty ends up far stronger than its size. Liu et al. (2025) analyse biases like this in GRPO; the lesson for anyone training or budgeting a reasoning model is that the exact shape of the reward decides how long the model thinks.
Chapter 6
Where reasoning still fails, and how to budget it
Everyday picture A long line of dominoes falls all the way only if every single one is placed right. Add more dominoes and the chance that one is misplaced grows, however careful you are with each.
Slips compound
Tiny example If each step of a chain slips 2% of the time, a 10-step chain is clean 82% of the time, and a 50-step chain only 36% of the time.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| chance one step slips | 0.02 | |
| chance one step is right | 0.98 | |
| steps in the chain | 50 | |
| chance all are right, multiplying because each step is a separate chance to slip | 0.364 |
In words: "the chance a whole chain is clean is the chance one step is right, multiplied by itself once per step."
With the numbers: , , .
Level 3: in Python
eps = 0.02
round((1 - eps) ** 10, 3) # → 0.817
round((1 - eps) ** 50, 3) # → 0.364
round((1 - eps) ** 200, 3) # → 0.018
Figure 13 · Drawn from the lesson's code
At a 2% slip rate the chance of a clean chain falls to 0.36 by 50 steps and near zero by 200; at 0.5% it falls much more slowly
primer.agents.planning.In code: steps_all_right is the formula.
What thinking costs
Tiny example Eight sampled chains of 2,000 tokens each, at an illustrative $10 per million output tokens, cost 16 cents per question. Run side by side at 50 tokens per second, the user waits 40 seconds whether you sample one chain or eight.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| chains sampled for one question | 8 | |
| tokens in each chain | 2,000 | |
| price per million output tokens | $10 | |
| one million: turns a per-million price into a per-token one | ||
| generation speed, tokens per second | 50 |
In words: "money grows with every token of every chain; waiting time, with chains run side by side, grows only with the length of one chain."
With the numbers: 8 × 2,000 × 10 / 1,000,000 = $0.16, and 2,000 / 50 = 40 seconds. One chain instead of eight costs $0.02 and still takes 40 seconds.
Level 3: in Python
n, T, c, v = 8, 2000, 10.0, 50
round(n * T * c / 10**6, 2) # → 0.16
T / v # → 40.0
# one chain instead of eight: an eighth of the money, the same wait
round(1 * T * c / 10**6, 2) # → 0.02
Thinking wider costs money; thinking longer costs money and time. See
primer.agents.cost for pricing requests and measuring cost per successful
task.
In code: reasoning_cost returns both numbers for one question.
Other ways reasoning fails
- The chain is not a transcript. The written reasoning need not be the real cause of the answer. Turpin et al. (2023) nudged models towards an answer with a hidden bias in the prompt; the answers followed the bias, and the written explanations never mentioned it. Treat a chain of thought as evidence about the model's reasoning, not a faithful record of it.
- Overthinking. Long chains on easy questions waste tokens and time, and a model can talk itself out of a right first answer.
- No checker, less progress. Reinforcement learning with verifiable rewards works where a program can check the answer. Open-ended writing, judgement calls and long projects have no cheap checker, and gains there are smaller.
- Checkers get gamed. A verifier with gaps is a target: code that
special-cases the unit tests passes them. This is reward hacking; see
primer.ml.reinforcement.
Budgeting it
Everyday picture You wouldn't convene a committee to decide what to have for lunch, and you wouldn't let one person sign off a bridge design alone. Match the effort to the stakes and to whether the result can be checked.
Figure 14 · Diagram
flowchart TD
Q[Incoming task] --> E{"Easy, or<br/>latency-critical?"}
E -->|yes| N["No thinking,<br/>or a small budget"]
E -->|no| C{"Can a program<br/>check the answer?"}
C -->|"yes: tests, math"| BV["Think, sample n,<br/>keep what passes the check"]
C -->|no| B["Think with a capped budget;<br/>vote if answers compare exactly"]
N & BV & B --> M["Measure accuracy and cost per task;<br/>raise budgets only where it pays"]
primer.agents.planning for decomposing long tasks into checkable steps.Test yourself
8 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1How can writing its reasoning out make a model more accurate, when its weights don't change?Think it through, then reveal
Each token is one forward pass with a fixed amount of serial computation. A problem that needs more sequential steps than one pass can do can't be solved in a single token. Writing intermediate results spreads the work over many passes, and the written text carries each result to the next pass: it is the model's working memory.
Question 2What is test-time compute, and what are the two basic ways to spend it?Think it through, then reveal
Computation spent while answering rather than while training. Think longer (one chain with a bigger thinking budget, which costs time and money) or think wider (many chains in parallel, then vote or verify, which costs money but no extra waiting when run side by side).
Question 3Why does majority voting over samples help, and when does it stop helping?Think it through, then reveal
If each sample is independently right more often than any single wrong answer appears, the right answer becomes the biggest pile as samples grow. It stops helping when samples share the same mistakes (they are one opinion repeated), when the model is usually wrong in one consistent way, or when answers can't be compared exactly.
Question 4What is the difference between an outcome reward model and a process reward model?Think it through, then reveal
An outcome reward model scores only the final answer, so it can't tell a lucky answer from a sound one and can't say where a wrong chain went wrong. A process reward model scores each step, catching errors that cancel out and pointing at the first bad step.
Question 5What is pass@n, and why is it a ceiling for best-of-n?Think it through, then reveal
The chance that at least one of n samples is right, 1 − (1 − p)ⁿ. A best-of-n picker can only choose among the samples it has, so even a perfect verifier can't do better than "some sample was right".
Question 6How does reinforcement learning with verifiable rewards teach a model to reason, if nobody grades the reasoning?Think it through, then reveal
Each problem gets a group of sampled attempts, each rewarded 1 or 0 by a program that checks the final answer. Attempts better than their group's average are made more likely, worse ones less likely. Whatever the right attempts had in common, including longer chains and rechecking, is reinforced. A group where every attempt scores the same teaches nothing, so the useful problems are ones the model solves only sometimes.
Question 7Why do reasoning models sometimes spend thousands of tokens on easy questions, and what reins that in?Think it through, then reveal
A reward for correctness alone never says stop: any extra length that helps slightly is reinforced. A per-token penalty in training, a thinking budget at answer time, and routing easy tasks to little or no thinking all rein it in. The exact shape of the reward matters: dividing by the group's spread, as GRPO does, can magnify a tiny length penalty.
Question 8Is a model's chain of thought an accurate account of why it answered?Think it through, then reveal
Not necessarily. Experiments that plant a hidden bias in a prompt show answers following the bias while the written reasoning never mentions it. The chain is useful evidence and often helpful, but it is not a guaranteed record of the computation.
Primary sources
The papers behind this lesson
Showed that prompting with a few worked examples that include intermediate steps makes large models far better at arithmetic, commonsense and symbolic reasoning.
Read the annotated companion →The paper ↗Introduced sampling many chains of thought and taking a majority vote over their final answers.
Read the annotated companion →The paper ↗Introduced the GSM8K dataset and showed that a trained verifier picking the best of many sampled solutions beats fine-tuning alone.
Read the annotated companion →The paper ↗Showed that process supervision, a reward model trained on labels for each step, selects correct solutions more reliably than outcome supervision.
Read the annotated companion →The paper ↗Measured how to spend test-time compute (longer revisions or verifier-guided search) and showed that adapting it to each prompt's difficulty pays.
The paper ↗Introduced GRPO, reinforcement learning that uses a group of sampled answers as its own baseline.
Read the annotated companion →The paper ↗Showed that reinforcement learning with verifiable rewards alone makes long, self-checking chains of thought emerge.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Wei et al., Chain-of-Thought Prompting (2022): https://arxiv.org/abs/2201.11903
- Kojima et al., Large Language Models are Zero-Shot Reasoners ("let's think step by step", 2022): https://arxiv.org/abs/2205.11916
- Nye et al., Show Your Work: Scratchpads for Intermediate Computation with Language Models (2021): https://arxiv.org/abs/2112.00114
- Li et al., Chain of Thought Empowers Transformers to Solve Inherently Serial Problems (2024): https://arxiv.org/abs/2402.12875
- Wang et al., Self-Consistency (2022): https://arxiv.org/abs/2203.11171
- Brown et al., Large Language Monkeys: Scaling Inference Compute with Repeated Sampling (2024): https://arxiv.org/abs/2407.21787
- Cobbe et al., Training Verifiers to Solve Math Word Problems (2021): https://arxiv.org/abs/2110.14168
- Uesato et al., Solving math word problems with process- and outcome-based feedback (2022): https://arxiv.org/abs/2211.14275
- Lightman et al., Let's Verify Step by Step (2023): https://arxiv.org/abs/2305.20050
- Snell et al., Scaling LLM Test-Time Compute Optimally (2024): https://arxiv.org/abs/2408.03314
- Muennighoff et al., s1: Simple test-time scaling (2025): https://arxiv.org/abs/2501.19393
- Zelikman et al., STaR: Bootstrapping Reasoning With Reasoning (2022): https://arxiv.org/abs/2203.14465
- Shao et al., DeepSeekMath (GRPO, 2024): https://arxiv.org/abs/2402.03300
- DeepSeek-AI, DeepSeek-R1 (2025): https://arxiv.org/abs/2501.12948
- Liu et al., Understanding R1-Zero-Like Training: A Critical Perspective (2025): https://arxiv.org/abs/2503.20783
- Chen et al., Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs (2024): https://arxiv.org/abs/2412.21187
- Turpin et al., Language Models Don't Always Say What They Think (2023): https://arxiv.org/abs/2305.04388
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.