rumblr Work in progressWIP

● The AI Primer · Lesson 13 · Part 1: how the model works inside

Reasoning models

thinking before answering

This lesson covers Chain of thought, test-time compute, verifiers, learning to reason with RL

Members · open during launch 47 min14 figures and diagrams9 interactive
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Chain of thought: every written token is another forward pass, so writing steps gives a fixed-depth model more serial computation, and the text is its working memory.
  2. Test-time compute: spend more tokens at answer time, either one longer chain or many chains; accuracy often grows with the logarithm of the budget until the problems run out.
  3. Self-consistency: sample several chains and vote on the final answer; it helps when the right answer is the most common one and mistakes are independent, and stalls when they are shared.
  4. Verifiers: outcome checks score the answer, process checks score each step; best-of-n with a good checker approaches pass@n = 1 − (1 − p)ⁿ.
  5. RL with verifiable rewards: reward right final answers, compare each attempt with its group (GRPO), and longer, self-checking reasoning emerges because it pays; a length penalty keeps it from overthinking.
  6. Cost and limits: thinking is paid per token, slips compound over long chains, and the written chain is not a guaranteed account of why the model answered.

Level 2

How it works, from scratch

Ask someone to multiply 37 by 48 in their head, instantly, and they will probably guess. Hand them a pencil and a minute and they will get it right. They didn't get smarter in that minute. They got room: somewhere to write 37 × 8 = 296 and 37 × 40 = 1,480, so each small step could lean on the last.

A language model is in the same position. It can do a fixed amount of work for each token it writes: one pass through all of its layers. A reasoning model is a model that has learned to use the pencil. Before it answers, it writes out intermediate steps, checks them, backs up when one is wrong, and only then commits to an answer. This lesson builds each piece from scratch:

  1. why writing steps helps at all (chain of thought);
  2. why more thinking can buy more accuracy (test-time compute);
  3. sampling several attempts and voting on the answer (self-consistency);
  4. checking attempts, either the final answer or every step (verifiers);
  5. how reinforcement learning teaches a model to think this way;
  6. what thinking costs, and where it still fails.

The pencil from the first paragraph, in your hands.

Chapter 1

Why thinking out loud helps

Run the clerk yourself first; the table below will then read as what it is.

Everyday picture Picture a clerk who, each time they glance at a page, can do exactly one addition and write down one number. Hand them a column of four numbers and allow one glance, and they have to guess. Allow three glances and a margin to write in, and they get it right every time. The margin is the whole trick: what they write on one glance is there to read on the next.

Tiny example Our toy model is that clerk. Every token it writes costs one forward pass (one run of the whole model to produce one token), and each forward pass can do one addition. Ask it for 3 + 5 + 8 + 2, which needs three additions.

Pass Answer at once (1 token) Think first, then answer
1 adds 3 + 5 = 8, has no pass left for 8 and 2, estimates them at 4.5 each: 8 + 9 = 17 ✗ writes 3+5=8
2 reads 8, writes 8+8=16
3 reads 16, answers 16 + 2 = 18 ✓

The weights are identical in both columns. The only difference is that the second column wrote its running totals down, where the next pass could read them. Writing intermediate steps before the answer is called chain of thought.

Figure 1 · Diagram

Reading it: each box is one forward pass, and each arrow is text handed to the next pass. In the top row there is only one box, so all the work has to fit in it, and it doesn't. In the bottom row every box does one small step and leaves its result in the text. Nothing inside the model carries a running total from one token to the next except what has been written, so the written steps are the model's working memory.

The rule behind the table: the number of steps a model can do one after another grows with the number of tokens it writes.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
serial steps: how many steps the model can chain, each using the result of the one before 1 or 3
serial steps one forward pass can do; for a transformer, roughly its number of layers 1 addition in the toy
tokens generated, the answer included; each token is one forward pass 1 at once, 3 thinking first
multiply

In words: "the length of the chain of steps a model can work through equals the steps per token times the number of tokens it writes."

With the numbers: the toy has and the sum needs 3 additions. Answering at once, , so : two additions short, so it guesses. Thinking first, , so : exactly enough. A 32-layer model answering at once has at most 32 layers of one-after-another work; with a 1,000-token chain of thought it has up to 32,000.

Level 3: in Python
# the toy: 1 addition per pass; 3 + 5 + 8 + 2 needs 3 additions
L = 1
# answering at once: T = 1, so S = 1, not enough
L * 1  # → 1
# the one pass adds 3 + 5, then estimates the unreached 8 and 2 at 4.5 each
(3 + 5) + round(4.5 * 2)  # → 17
# thinking first: T = 3, so S = 3, exactly enough
L * 3  # → 3
t1 = 3 + 5  # → 8
t2 = t1 + 8  # → 16
t2 + 2  # → 18
# a 32-layer model: answering at once, then after a 1,000-token chain of thought
32 * 1  # → 32
32 * 1000  # → 32000

A layer is not literally one addition, so is a rough count, but the shape of the rule holds: a model with a fixed depth can only do a fixed amount of one-after-another work per token, and problems that need more must spread it over more tokens. Li et al. (2024) proved a version of this for transformers: with enough chain-of-thought tokens they can solve inherently serial problems that a fixed-depth transformer answering at once cannot.

Figure 2 · Drawn from the lesson's code

2 4 6 8 10 additions the problem needs 0.0 0.2 0.4 0.6 0.8 1.0 accuracy One addition per pass: steps are the only way answer in 1 token write steps, then answer 2 4 6 8 10 additions the problem needs 2 4 6 8 10 tokens (forward passes) used The price: one token per step

Answering in one token, the toy is right only on one-addition sums and about 5% of the rest; writing steps, it is always right, at one token per addition

Reading it: on the left, the x-axis is how many additions a sum needs and the y-axis is how often the toy gets it right. Answering in one token (red), it is perfect on one-addition sums and then collapses to the few percent of lucky estimates. Writing steps (blue), it is right every time. On the right is the price: the blue line climbs one token per addition, while the red line stays at one. Chain of thought doesn't make the model smarter per token; it lets the model buy more tokens.

That price is real: each token costs roughly floating-point operations for a model with parameters (see primer.ml.inference), so a 1,000-token chain costs a thousand answers' worth of compute. In practice this is why simply asking a model to "think step by step" (Kojima et al., 2022) or showing it worked examples with steps (Wei et al., 2022) improved accuracy on arithmetic and logic puzzles, and why every reasoning model writes a long chain before its answer.

In code: solve is the toy model: it spends one forward pass per written step, one on the answer, and estimates anything it never reached. It returns an Attempt holding the steps, the answer and the tokens used.

Chapter 2

Test-time compute: buying accuracy with tokens

Everyday picture An exam has questions of very different difficulty. Give yourself one minute per question and you finish only the easiest. Two minutes, and the next tier falls. Each doubling of time unlocks one more tier, until you can finish everything and extra time buys nothing.

Tiny example Test-time compute is computation spent while answering, as opposed to while training. The simplest dial is a thinking budget: the most tokens the model may write before it must answer. Take six problems that need 1, 2, 4, 8, 16 and 32 additions, and a budget of 8 tokens. Four of them fit (1, 2, 4, 8). The other two must be guessed, and say a guess is right 10% of the time. Accuracy is 4/6 + 2/6 × 0.1 = 0.70.

There are two ways to spend test-time compute, and the rest of this lesson uses both.

Figure 3 · Diagram

Reading it: the top path spends compute sequentially: one chain, allowed to run longer. It helps when the problem needs many steps in a row. The bottom path spends it in parallel: several independent chains, then a rule that picks one answer. It helps when the model is right sometimes but not reliably. The top path costs waiting time; the bottom path costs money but, run side by side, no extra waiting. Sections 3 and 4 are about the "Pick one" box.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the thinking budget, in tokens 8
the share of problems that need at most steps, so fit in the budget 4/6 = 0.667
the share that doesn't fit and must be guessed 2/6 = 0.333
the chance a guess happens to be right 0.1
the share of problems answered correctly with budget 0.70

In words: "accuracy is the share of problems that fit in the budget, plus a lucky share of the ones that don't."

With the numbers: , so . Doubling to lets a fifth level fit: , . Every doubling adds the same , until fits everything and .

Level 3: in Python
needs = [1, 2, 4, 8, 16, 32]
g = 0.1
B = 8
# F(B): the share of problems that fit
F = sum(d <= B for d in needs) / len(needs)
round(F, 3)  # → 0.667
# acc(B) = F(B) + (1 − F(B)) · g
round(F + (1 - F) * g, 3)  # → 0.7
# each doubling of B lets one more of the six levels fit
round((1 / len(needs)) * (1 - g), 3)  # → 0.15

Figure 4 · Drawn from the lesson's code

2 0 2 1 2 2 2 3 2 4 2 5 2 6 thinking budget B (tokens, log scale) 0.0 0.2 0.4 0.6 0.8 1.0 accuracy Each doubling of budget buys the same step formula, g = 0.1 simulated sums 2 0 2 1 2 2 2 3 2 4 2 5 2 6 thinking budget B (tokens, log scale) 2 0 2 1 2 2 2 3 2 4 2 5 2 6 tokens actually spent (mean) Easy problems stop early; past 32, nothing changes budget allowed

Accuracy climbs 0.15 per doubling of the budget, from 0.25 at 1 token to 1.0 at 32, and the simulated sums land on the formula; tokens actually spent level off at about 10

Reading it: on the left, the x-axis is the budget on a doubling (log) scale. The grey line is the formula; the blue dots are the toy model from section 1 solving real random sums. On a log axis a straight line means "each doubling adds the same amount", and that is what both show until , where every problem fits and the line goes flat. The dots sit a little below the line at small budgets because a guess about many unreached numbers is right less often than 10%. On the right is what was actually spent: always less than the budget, because easy problems stop early, and nothing more past 32.

The straight line on a log axis is built into this toy, because its difficulties are spaced by doubling. Real problem sets are also spread over many scales of difficulty, and published reasoning models show the same shape over a useful range: accuracy rising roughly in proportion to the logarithm of thinking tokens, then leveling off. Snell et al. (2024) found that spending test-time compute adaptively (more on harder prompts) beats spending it evenly, and that on problems a small model can sometimes solve, extra test-time compute can stand in for a much larger model. In practice, model APIs expose this dial as a "reasoning effort" setting or a maximum number of thinking tokens.

In code: budget_accuracy is the formula, and budget_sweep runs solve on random sums at each budget and reports accuracy and tokens actually spent.

Chapter 3

Sampling many and voting: self-consistency

Everyday picture Unsure of an answer, you ask five friends separately. If four of them say the same thing, you trust it. Each friend can be wrong, but it is unlikely that most of them are wrong in the same way, unless they all read the same wrong article.

Tiny example A model that samples its tokens (see primer.ml.inference) writes a different chain each time. Ask the toy's noisy cousin for 3 + 5 + 8 + 2 three times and it answers 18, 17, 18. Keep only the final answers and take the most common one: 18. Sampling several chains of thought and returning the most common final answer is self-consistency (Wang et al., 2022), also called majority voting.

Figure 5 · Diagram

Reading it: the question fans out into independent chains, each free to take a different route. The chains themselves are thrown away; only their last lines meet in the counting box. That is why voting needs answers that can be compared exactly (a number, a multiple-choice letter): two essays are never identical, so there would be nothing to count.

How much does voting help? Take the simplest case, a yes/no question: every wrong vote lands on the same wrong answer, and each vote is right with the same chance , independently of the others.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
number of sampled votes (odd, so there are no ties) 3
chance one vote is right 0.6
how many of the votes are right 2 or 3
rounded up: the smallest number of votes that wins 2
" choose ": how many ways to pick which of the votes are the right ones
the chance of one particular pattern: those right, the other wrong
add up over every winning count of right votes
the chance the majority is right 0.648

In words: "the vote is right when at least half the votes are right, so add up the chance of every such count: the number of ways to get that many right, times the chance of each way."

With the numbers: , . Two right: $3 \times 0.6^2 \times 0.4 = 0.4320.6^3 = 0.216$. Together, 0.648, up from 0.6 for one vote. Five votes give 0.683. But a solver right only 40% of the time gets worse with five votes: 0.317. Voting amplifies whichever answer is most common, right or wrong. (This is the Condorcet jury theorem, from 1785.)

Level 3: in Python
import math
p, n = 0.6, 3
# one term per winning count k = 2, 3
terms = [math.comb(n, k) * p**k * (1 - p)**(n - k) for k in range(math.ceil(n / 2), n + 1)]
[round(t, 3) for t in terms]  # → [0.432, 0.216]
round(sum(terms), 3)  # → 0.648
# five votes
round(sum(math.comb(5, k) * 0.6**k * 0.4**(5 - k) for k in range(3, 6)), 3)  # → 0.683
# a 40% solver on a yes/no question: voting makes it worse
round(sum(math.comb(5, k) * 0.4**k * 0.6**(5 - k) for k in range(3, 6)), 3)  # → 0.317

Figure 6 · Drawn from the lesson's code

0 5 10 15 20 25 30 samples voting, n 0.0 0.2 0.4 0.6 0.8 1.0 accuracy of the vote Voting amplifies whatever answer is most common yes/no question, p = 0.8 yes/no question, p = 0.6 yes/no question, p = 0.4 p = 0.4, wrong answers scattered over 10 values

Votes of an 80% or 60% solver climb towards 1, a 40% solver on a yes/no question sinks towards 0, but the same 40% solver climbs past 0.9 when its wrong answers scatter over ten values

Reading it: the x-axis is how many samples vote; the y-axis is how often the vote is right. The solid lines are the formula. Above 0.5 (green, blue) more votes push accuracy towards 1; below 0.5 (solid red) they push it towards 0. The dashed red line is the same 40% solver on a question with a numeric answer, where its mistakes scatter over ten different wrong values. No wrong value gets more than a few percent of the votes, so 40% is the biggest pile and the vote climbs past 0.9. That is the situation self-consistency relies on in math problems: the right answer only needs to be the most common one, not a majority.

When voting fails: shared mistakes. Everything above assumed the votes were independent. Samples from one model are not: they share its training, its blind spots and its reading of the question. Model that as a shared draw: each sample copies one common answer with probability (rho), and otherwise answers on its own. Each single sample is still right 60% of the time, so only the correlation changes.

Figure 7 · Drawn from the lesson's code

0 5 10 15 20 25 30 samples voting, n 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 accuracy of the vote Same per-sample accuracy, very different votes one sample alone: 0.6 shared-mistake rate rho = 0.0 shared-mistake rate rho = 0.3 shared-mistake rate rho = 0.6 shared-mistake rate rho = 0.9

With independent samples 31 votes reach nearly 1.0; with rho = 0.3 the vote levels off near 0.86; with rho = 0.6 or 0.9 it stays at about 0.6, no better than one sample

Reading it: every line starts at 0.6 on the left, because one sample is one sample. Independent samples (blue) climb to nearly 1. At (amber) and (red) the lines go flat at 0.6: whenever the shared answer is wrong, the 60% or 90% of votes that copy it outnumber the at most 40% × 0.6 = 24% (or 6%) that can independently land on the right answer. Thirty-one correlated votes are one opinion, repeated. At (green) the right answer's share, 0.7 × 0.6 = 42%, still beats the shared wrong answer's 0.3 + 0.7 × 0.4 / 5 = 35.6%, so the vote does win in the long run, just slowly.

In practice, self-consistency gives a solid gain for a few samples and then flattens, because a model's mistakes on a question are correlated. Diversity helps (different prompts, different models, a tool that computes instead of guessing), and voting only works where answers can be compared exactly.

In code: majority_vote picks the most common answer, majority_accuracy is the formula (with ties on even counted as a coin flip), and correlated_vote_accuracy simulates votes with scattered wrong answers and a shared-mistake rate.

The binomial sum above, as bars you can drag.

Chapter 4

Verifiers: checking the answer, or checking every step

Plant the slips yourself before reading how the two checkers treat them.

Everyday picture One teacher looks only at the boxed answer at the bottom of the page. Another marks every line of working. The first can't tell a lucky guess from understanding, and when the answer is wrong, can't tell you where you went wrong. The second can do both.

Tiny example A verifier is anything that scores a candidate solution. Here are two chains the toy wrote for 3 + 5 + 8 + 2 (true answer 18), each with slips:

Chain Final answer Outcome check: is it 18? Step check: first line that is false
3+5=8, 8+8=17, 17+2=19 19 reward 0 step 2: 8 + 8 is 16
3+5=9, 9+8=16, 16+2=18 18 reward 1 step 1: 3 + 5 is 8

An outcome verifier looks only at the final answer: either a learned outcome reward model (ORM) that predicts whether it is right, or a check against a known answer. It gives the second chain full marks, although two slips just happened to cancel. A process verifier, or process reward model (PRM), scores every step. It catches the second chain's slip and points at exactly where the first one went wrong.

Figure 8 · Diagram

Reading it: both rows read the same chain. The outcome verifier jumps straight to the diamond at the end and returns one bit. The process verifier stamps every box. Notice that step 3, 17 + 2 = 19, is marked ok: it is correct arithmetic on a wrong input. A step checker judges each step on its own terms, and the first bad step is where the chain left the rails.

Verifiers power best-of-n: sample chains, keep the one the verifier scores highest. With a perfect verifier, best-of-n is right whenever any of the samples is right, a number called pass@n.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
chance one sample is right 0.3
number of samples drawn 5
chance one sample is wrong 0.7
chance all are wrong, multiplying because samples are independent
chance at least one of the is right 0.832

In words: "the chance that at least one sample is right is one minus the chance that every sample is wrong."

With the numbers: a model right 30% of the time is wrong on five tries in a row with chance , so pass@5 = 0.832. A 30% model with a perfect checker and five tries beats an 80% model on one try. In the experiment below, each chain has 5 additions that each slip with chance 0.2, so a chain is clean with chance .

Level 3: in Python
p, n = 0.3, 5
# (1 − p)^n: every one of the five is wrong
round((1 - p) ** n, 5)  # → 0.16807
round(1 - (1 - p) ** n, 5)  # → 0.83193
# the experiment's chains: 5 additions, each clean with chance 0.8
round(0.8 ** 5, 3)  # → 0.328

Figure 9 · Drawn from the lesson's code

2 0 2 1 2 2 2 3 2 4 2 5 chains sampled, n (log scale) 0.0 0.2 0.4 0.6 0.8 1.0 accuracy A good checker turns many tries into a right answer pass@n: any sample right (ceiling) best-of-n, step checker picks majority vote one sample

With 8 chains, one sample is right 38% of the time, a majority vote 67%, best-of-8 with a step checker 96%, against a pass@8 ceiling of 99%

Reading it: the x-axis doubles the number of chains sampled; the y-axis is accuracy. One sample (red) stays at 0.38 whatever is: about a third of chains are clean, and a few more land on 18 by luck. Voting (blue) climbs slowly, because wrong answers bunch on near misses one or two away from the truth. Best-of-n with the step checker (green) hugs the dotted pass@n ceiling: at 8 chains, 0.96 against 0.99. The small gap is the lucky chains: their slips cancelled, so pass@n counts them as right, but the checker refuses them. That is the checker doing its job.

In practice, the best verifiers are programs: unit tests for code, an exact match for a math answer, a proof checker. Where no program can check, a learned verifier stands in. Lightman et al. (2023) trained a process reward model on 800,000 human labels of individual steps and found it picked correct solutions far more reliably than an outcome reward model. A learned verifier can be fooled, though, and more samples mean more chances to find a wrong answer it likes: Cobbe et al. (2021) saw best-of-n accuracy start to fall again after a few hundred samples.

In code: check_step is the toy's process verifier for one line, first_bad_step runs it over a chain, and outcome_reward compares only the final answer with a reference. noisy_chain writes chains whose slips carry forward, pass_at_n is the formula, and verifier_experiment compares one sample, voting, best-of-n with the step checker, and the pass@n ceiling.

Chapter 5

Learning to reason with reinforcement learning

Flip the attempts in one group first; the advantage formula will then read as what it is.

Everyday picture A workbook with the answers printed in the back. Nobody shows you how to solve anything. You try each problem several ways, check the back, and do more of whatever worked. Over hundreds of problems you discover good habits on your own: write things down, double-check the tricky step, don't stop too early.

Tiny example Prompting a model to "think step by step" gets it to write steps; a reasoning model is trained to write good ones. The recipe behind models such as DeepSeek-R1 is reinforcement learning (see primer.ml.reinforcement) with two ingredients:

  • Verifiable rewards. Train on problems whose answers a program can check: math with a known final answer, code with unit tests. The reward is 1 if the final answer is right and 0 if not. Nobody grades the chain itself.
  • Group comparisons. For each problem, sample a group of attempts and judge each one against its own group. With and rewards 1, 0, 0, 1, the group averages 0.5, so the two right attempts are above average (+1) and the two wrong ones below (−1). The model is nudged towards whatever the right ones did, including how they reasoned. This is the heart of GRPO (group relative policy optimization).

Figure 10 · Diagram

Reading it: follow the loop clockwise. Nothing in it ever says "think longer" or "check your work"; the only signal is whether the last line was right. Whatever habits the chains of right attempts share get reinforced, batch after batch. The group is what makes this work without a separate model estimating how good an attempt "should" be: the other attempts at the same problem are the baseline.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
how many attempts are sampled for one problem 4
which attempt, from 1 to 1, 2, 3, 4
attempt 's reward: 1 if its final answer is right, 0 if not 1, 0, 0, 1
the average reward in the group 0.5
the standard deviation: the typical distance of a reward from the mean (see primer.notation) 0.5
attempt 's advantage: how much better than its group it did, in units of the group's spread +1, −1, −1, +1

In words: "an attempt's advantage is how far its reward sits above its group's average, measured in units of how spread out the group's rewards are."

With the numbers: rewards 1, 0, 0, 1 have mean 0.5 and standard deviation 0.5, so the advantages are (1 − 0.5)/0.5 = +1 and (0 − 0.5)/0.5 = −1. If all four attempts are right, every reward equals the mean and the spread is 0: every advantage is 0, and the problem teaches nothing. The same holds when all four are wrong. Training therefore needs problems the model solves sometimes.

Level 3: in Python
import statistics
r = [1, 0, 0, 1]
mu = statistics.mean(r)  # → 0.5
sigma = statistics.pstdev(r)  # → 0.5
[(r_i - mu) / sigma for r_i in r]  # → [1.0, -1.0, -1.0, 1.0]
# a group that all succeeded: spread 0, nothing to learn
statistics.pstdev([1, 1, 1, 1])  # → 0.0

The toy: learning how long to think. Our policy (the model's rule for choosing what to do) makes one choice per problem: how many tokens to think for, out of 1, 2, 4, 8, 16 or 32. It keeps a separate choice for each difficulty (problems needing 1, 2, 4 or 8 steps), and it starts out preferring short answers, like a model never rewarded for thinking: 63% of the time it answers in 1 token. A chain shorter than the steps needed must guess (right 10% of the time). Otherwise each step gets as many tries as the length allows, and a try slips 10% of the time: one pass through the steps, or two (the first attempt and a recheck that catches a slip), or more. The reward is 1 for a right answer and 0 for a wrong one. That is all.

Figure 11 · Drawn from the lesson's code

0 100 200 300 400 500 600 training iteration 2 3 4 5 6 7 mean thinking length (tokens) It learns to write longer 0 100 200 300 400 500 600 training iteration 0.4 0.5 0.6 0.7 0.8 0.9 1.0 accuracy ...and gets more right correctness only penalty 0.01 per token, ÷ spread (GRPO) penalty 0.01 per token, no ÷ spread

Rewarded only for correct answers, mean thinking length rises from 2 to about 7.6 tokens and accuracy from 0.41 to 0.96; with a length penalty the length and accuracy settle a little lower

Reading it: the x-axis is training time. On the left, the blue line (reward for correctness only) shows the average thinking length rising from 2 tokens to about 7.6; on the right, accuracy rising from 0.41 to 0.96. Nobody asked for longer answers: longer answers were simply right more often, so they were reinforced. This is the toy version of what DeepSeek-R1 reported at scale: the length of its chains grew steadily through reinforcement learning, and behaviours like re-checking and backing up appeared without being taught. The red and green lines add a price per token; they come next.

Figure 12 · Drawn from the lesson's code

d = 1 d = 2 d = 4 d = 8 difficulty: additions the problem needs 0 2 4 6 8 10 12 14 16 18 tokens the trained model spends Hard problems get more thinking, with room to recheck steps needed d 2d: room for one recheck correctness only penalty 0.01 per token, ÷ spread (GRPO) penalty 0.01 per token, no ÷ spread

Trained on correctness alone, the model spends about 2.5, 4, 8 and 16 tokens on problems needing 1, 2, 4 and 8 steps: at or just past twice the steps needed

Reading it: each group of bars is one difficulty; bar height is how many tokens the trained model spends on it. The black tick is the steps the problem needs, the grey tick twice that. The blue bars (correctness only) sit on the grey ticks: the model learned to spend more on harder problems, and to leave room for one recheck of every step, which lifts the hardest problems from 0.43 to 0.92. That extra room is self-correction in miniature: the chain gets longer because checking pays.

In code: group_advantages is the formula, success_probability is the toy's chance of solving a problem at a given length, and train_reasoner runs the whole loop: sample a group of lengths per difficulty, reward, compute advantages, and nudge the policy.

The price of thinking

Rewarded for correctness alone, nothing ever tells the model to stop: any extra length that helps even slightly gets reinforced. Real reasoning models show this as overthinking: hundreds of tokens spent on "what is 2 + 3?". The usual remedy is to charge for length: subtract a small penalty per token from the reward. The expected reward for a chain of length on a problem needing steps (with ) becomes:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the expected (average) reward for this length 0.763
steps the problem needs 8
thinking length, in tokens 16
rounded down: how many tries each step gets 2
chance one try at a step slips 0.1
chance every try at one step slips, so the step fails 0.01
chance all steps come out right 0.923
lambda: the penalty per token of thinking 0.01
the total charge for the chain 0.16

In words: "the reward for a length is how likely it is to get every step right, given the tries it allows, minus a small charge for every token."

With the numbers: for an 8-step problem, 8 tokens give ; 16 tokens give ; 32 tokens give . Sixteen wins: one recheck is worth paying for, a third is not. For a 1-step problem, 1 token gives , 2 tokens , 4 tokens . Two wins.

Level 3: in Python
eps, lam = 0.1, 0.01
# an 8-step problem: chance of success at 8, 16 and 32 tokens
[round((1 - eps ** (L // 8)) ** 8, 3) for L in (8, 16, 32)]  # → [0.43, 0.923, 0.999]
# minus λL: 16 tokens wins
[round((1 - eps ** (L // 8)) ** 8 - lam * L, 3) for L in (8, 16, 32)]  # → [0.35, 0.763, 0.679]
# a 1-step problem at 1, 2 and 4 tokens: 2 tokens wins
[round((1 - eps ** (L // 1)) ** 1 - lam * L, 3) for L in (1, 2, 4)]  # → [0.89, 0.97, 0.96]

Now look back at the two figures. With the penalty and no division by the spread (green), training finds exactly those best lengths, 2 tokens for the easiest problems and 16 for the hardest. With GRPO's division by the spread (red), the easiest problems shrink to 1 token, and accuracy settles at 0.92, below the 0.96 of correctness alone. Why? In a group where every attempt succeeded, rewards differ only by the penalty: 0.99 for 1 token, 0.98 for 2. Their spread is tiny, so dividing by it inflates that 0.01 difference into advantages of +1 and −1, as loud as the difference between solved and failed. The penalty ends up far stronger than its size. Liu et al. (2025) analyse biases like this in GRPO; the lesson for anyone training or budgeting a reasoning model is that the exact shape of the reward decides how long the model thinks.

The expected-reward formula from "The price of thinking", with your own d, λ and ε in it.

Chapter 6

Where reasoning still fails, and how to budget it

Everyday picture A long line of dominoes falls all the way only if every single one is placed right. Add more dominoes and the chance that one is misplaced grows, however careful you are with each.

Slips compound

Tiny example If each step of a chain slips 2% of the time, a 10-step chain is clean 82% of the time, and a 50-step chain only 36% of the time.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
chance one step slips 0.02
chance one step is right 0.98
steps in the chain 50
chance all are right, multiplying because each step is a separate chance to slip 0.364

In words: "the chance a whole chain is clean is the chance one step is right, multiplied by itself once per step."

With the numbers: , , .

Level 3: in Python
eps = 0.02
round((1 - eps) ** 10, 3)  # → 0.817
round((1 - eps) ** 50, 3)  # → 0.364
round((1 - eps) ** 200, 3)  # → 0.018

Figure 13 · Drawn from the lesson's code

0 25 50 75 100 125 150 175 200 steps in the chain, k 0.0 0.2 0.4 0.6 0.8 1.0 chance every step is right Long chains need checking, not just length 50 steps at 2%: 0.36 slip rate per step = 0.005 slip rate per step = 0.02 slip rate per step = 0.05

At a 2% slip rate the chance of a clean chain falls to 0.36 by 50 steps and near zero by 200; at 0.5% it falls much more slowly

Reading it: the x-axis is the length of the chain; the y-axis is the chance it contains no slip at all. Every curve falls, and the only thing that flattens one is a lower slip rate per step. That is why length alone is not reasoning: the trained model in section 5 got better by spending its extra tokens on rechecking, which lowers the effective slip rate, and why verifiers that check each step matter. The same law governs agents taking many actions; see primer.agents.planning.

In code: steps_all_right is the formula.

What thinking costs

Tiny example Eight sampled chains of 2,000 tokens each, at an illustrative $10 per million output tokens, cost 16 cents per question. Run side by side at 50 tokens per second, the user waits 40 seconds whether you sample one chain or eight.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
chains sampled for one question 8
tokens in each chain 2,000
price per million output tokens $10
one million: turns a per-million price into a per-token one
generation speed, tokens per second 50

In words: "money grows with every token of every chain; waiting time, with chains run side by side, grows only with the length of one chain."

With the numbers: 8 × 2,000 × 10 / 1,000,000 = $0.16, and 2,000 / 50 = 40 seconds. One chain instead of eight costs $0.02 and still takes 40 seconds.

Level 3: in Python
n, T, c, v = 8, 2000, 10.0, 50
round(n * T * c / 10**6, 2)  # → 0.16
T / v  # → 40.0
# one chain instead of eight: an eighth of the money, the same wait
round(1 * T * c / 10**6, 2)  # → 0.02

Thinking wider costs money; thinking longer costs money and time. See primer.agents.cost for pricing requests and measuring cost per successful task.

In code: reasoning_cost returns both numbers for one question.

Other ways reasoning fails

  • The chain is not a transcript. The written reasoning need not be the real cause of the answer. Turpin et al. (2023) nudged models towards an answer with a hidden bias in the prompt; the answers followed the bias, and the written explanations never mentioned it. Treat a chain of thought as evidence about the model's reasoning, not a faithful record of it.
  • Overthinking. Long chains on easy questions waste tokens and time, and a model can talk itself out of a right first answer.
  • No checker, less progress. Reinforcement learning with verifiable rewards works where a program can check the answer. Open-ended writing, judgement calls and long projects have no cheap checker, and gains there are smaller.
  • Checkers get gamed. A verifier with gaps is a target: code that special-cases the unit tests passes them. This is reward hacking; see primer.ml.reinforcement.

Budgeting it

Everyday picture You wouldn't convene a committee to decide what to have for lunch, and you wouldn't let one person sign off a bridge design alone. Match the effort to the stakes and to whether the result can be checked.

Figure 14 · Diagram

Reading it: start at the top with each incoming task. The first question routes easy or time-critical work away from thinking entirely, because that is where overthinking wastes the most. The second asks whether a program can check the answer: if so, sampling several chains and keeping the one that passes turns spare compute into accuracy almost for free (section 4). If not, cap the budget and vote only when answers can be compared. Everything ends in the same box: measure, because the right budget is an empirical question per task, not a constant. See primer.agents.planning for decomposing long tasks into checkable steps.

The bill from "What thinking costs", with the sliders on it.

Test yourself

8 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1How can writing its reasoning out make a model more accurate, when its weights don't change?Think it through, then reveal

Each token is one forward pass with a fixed amount of serial computation. A problem that needs more sequential steps than one pass can do can't be solved in a single token. Writing intermediate results spreads the work over many passes, and the written text carries each result to the next pass: it is the model's working memory.

Question 2What is test-time compute, and what are the two basic ways to spend it?Think it through, then reveal

Computation spent while answering rather than while training. Think longer (one chain with a bigger thinking budget, which costs time and money) or think wider (many chains in parallel, then vote or verify, which costs money but no extra waiting when run side by side).

Question 3Why does majority voting over samples help, and when does it stop helping?Think it through, then reveal

If each sample is independently right more often than any single wrong answer appears, the right answer becomes the biggest pile as samples grow. It stops helping when samples share the same mistakes (they are one opinion repeated), when the model is usually wrong in one consistent way, or when answers can't be compared exactly.

Question 4What is the difference between an outcome reward model and a process reward model?Think it through, then reveal

An outcome reward model scores only the final answer, so it can't tell a lucky answer from a sound one and can't say where a wrong chain went wrong. A process reward model scores each step, catching errors that cancel out and pointing at the first bad step.

Question 5What is pass@n, and why is it a ceiling for best-of-n?Think it through, then reveal

The chance that at least one of n samples is right, 1 − (1 − p)ⁿ. A best-of-n picker can only choose among the samples it has, so even a perfect verifier can't do better than "some sample was right".

Question 6How does reinforcement learning with verifiable rewards teach a model to reason, if nobody grades the reasoning?Think it through, then reveal

Each problem gets a group of sampled attempts, each rewarded 1 or 0 by a program that checks the final answer. Attempts better than their group's average are made more likely, worse ones less likely. Whatever the right attempts had in common, including longer chains and rechecking, is reinforced. A group where every attempt scores the same teaches nothing, so the useful problems are ones the model solves only sometimes.

Question 7Why do reasoning models sometimes spend thousands of tokens on easy questions, and what reins that in?Think it through, then reveal

A reward for correctness alone never says stop: any extra length that helps slightly is reinforced. A per-token penalty in training, a thinking budget at answer time, and routing easy tasks to little or no thinking all rein it in. The exact shape of the reward matters: dividing by the group's spread, as GRPO does, can magnify a tiny length penalty.

Question 8Is a model's chain of thought an accurate account of why it answered?Think it through, then reveal

Not necessarily. Experiments that plant a hidden bias in a prompt show answers following the bias while the written reasoning never mentions it. The chain is useful evidence and often helpful, but it is not a guaranteed record of the computation.

Primary sources

The papers behind this lesson

Wei et al., Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (2022)

Showed that prompting with a few worked examples that include intermediate steps makes large models far better at arithmetic, commonsense and symbolic reasoning.

Read the annotated companion →The paper ↗
Wang et al., Self-Consistency Improves Chain of Thought Reasoning in Language Models (2022)

Introduced sampling many chains of thought and taking a majority vote over their final answers.

Read the annotated companion →The paper ↗
Cobbe et al., Training Verifiers to Solve Math Word Problems (2021)

Introduced the GSM8K dataset and showed that a trained verifier picking the best of many sampled solutions beats fine-tuning alone.

Read the annotated companion →The paper ↗
Lightman et al., Let's Verify Step by Step (2023)

Showed that process supervision, a reward model trained on labels for each step, selects correct solutions more reliably than outcome supervision.

Read the annotated companion →The paper ↗
Snell et al., Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters (2024)

Measured how to spend test-time compute (longer revisions or verifier-guided search) and showed that adapting it to each prompt's difficulty pays.

The paper ↗
Shao et al., DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models (2024)

Introduced GRPO, reinforcement learning that uses a group of sampled answers as its own baseline.

Read the annotated companion →The paper ↗
DeepSeek-AI, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (2025)

Showed that reinforcement learning with verifiable rewards alone makes long, self-checking chains of thought emerge.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Wei et al., Chain-of-Thought Prompting (2022): https://arxiv.org/abs/2201.11903
  • Kojima et al., Large Language Models are Zero-Shot Reasoners ("let's think step by step", 2022): https://arxiv.org/abs/2205.11916
  • Nye et al., Show Your Work: Scratchpads for Intermediate Computation with Language Models (2021): https://arxiv.org/abs/2112.00114
  • Li et al., Chain of Thought Empowers Transformers to Solve Inherently Serial Problems (2024): https://arxiv.org/abs/2402.12875
  • Wang et al., Self-Consistency (2022): https://arxiv.org/abs/2203.11171
  • Brown et al., Large Language Monkeys: Scaling Inference Compute with Repeated Sampling (2024): https://arxiv.org/abs/2407.21787
  • Cobbe et al., Training Verifiers to Solve Math Word Problems (2021): https://arxiv.org/abs/2110.14168
  • Uesato et al., Solving math word problems with process- and outcome-based feedback (2022): https://arxiv.org/abs/2211.14275
  • Lightman et al., Let's Verify Step by Step (2023): https://arxiv.org/abs/2305.20050
  • Snell et al., Scaling LLM Test-Time Compute Optimally (2024): https://arxiv.org/abs/2408.03314
  • Muennighoff et al., s1: Simple test-time scaling (2025): https://arxiv.org/abs/2501.19393
  • Zelikman et al., STaR: Bootstrapping Reasoning With Reasoning (2022): https://arxiv.org/abs/2203.14465
  • Shao et al., DeepSeekMath (GRPO, 2024): https://arxiv.org/abs/2402.03300
  • DeepSeek-AI, DeepSeek-R1 (2025): https://arxiv.org/abs/2501.12948
  • Liu et al., Understanding R1-Zero-Like Training: A Critical Perspective (2025): https://arxiv.org/abs/2503.20783
  • Chen et al., Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs (2024): https://arxiv.org/abs/2412.21187
  • Turpin et al., Language Models Don't Always Say What They Think (2023): https://arxiv.org/abs/2305.04388

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.