The lesson in one minute
What you'll be able to explain
- RL learns from a score, not an answer: sample an action from the policy, get a reward, make rewarded actions more likely.
- REINFORCE: step along reward × ∇ log π(action); for softmax, ∇ log π is one-hot minus the probabilities. Unbiased but noisy.
- Baselines subtract the typical reward, turning rewards into advantages. Same average gradient, far less variance.
- PPO reuses each batch for several steps, clips the probability ratio to 1 ± ε so no step goes too far, and uses a value network as baseline plus a KL leash to a reference model.
- GRPO drops the value network: sample a group of answers per prompt and normalise rewards within the group. With a verifier as reward, it is how reasoning models are trained.
- Reward hacking: the policy optimises the reward you wrote, not the goal you meant. Defend with verifiable rewards, better reward models, a KL leash and a held-out check of the real goal.
Level 1
The practitioner's guide
In one sentence
Reinforcement learning trains a model from a score on what it produced rather than from a correct answer to copy, which is how a language model learns things nobody can write down (be helpful, reason to the right answer) and how it learns to game a score that was written down badly.
When you need it
You need RL when you can judge an answer but cannot
write the perfect one: a proof either checks or it doesn't, tests pass or
fail, one reply is better than another though neither is "the" reply.
The tell: you find yourself writing a grader, a checker or a rubric instead
of example answers. You don't need it when you can write the answers
(supervised fine-tuning copies them, at a fraction of the cost) or when
you have pairs of better and worse answers and nothing more (DPO in
primer.ml.training_stages learns from pairs with no sampling loop). And
you never need to run the vendor's own RL: the helpfulness, refusals and
reasoning of a hosted model were trained this way before you arrived, which
is why this lesson matters even if you never train: it explains why models
answer at length, flatter, and sometimes optimise the letter of your
instruction instead of its spirit.
Your options
From the cheapest to the most committed:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Supervised fine-tuning on demonstrations | Copy correct answers you wrote | The behaviour in the examples, nothing beyond them | Writing the answers | Your training stack, or a hosted API |
| DPO on preference pairs | Learn from "this one beat that one", no sampling during training | Shifts tone and choices without a reward model or an RL loop | Thousands of comparisons | Your training stack, or a hosted API |
| Hosted reinforcement fine-tuning with your grader | The vendor samples answers and reinforces the ones your grader scores high | An RL loop you don't build; the grader is still yours to get right | Grader design, many sampled answers per prompt, the vendor's price | The vendor's API |
| GRPO with a verifiable reward | Sample a group of answers per prompt, check each, reinforce the above-average ones | A reward with no learned blind spot, and no second model to train | Generation dominates: 8 answers per prompt is a common default; a checker that cannot be argued with | Your training stack |
| PPO with a learned reward model | Train a reward model on ratings, then a value network and the policy against it, on a KL leash | Optimises a goal no program can check, such as helpfulness | Two extra models the size of the policy, and the reward model's blind spots to defend against | Your training stack |
How to choose
Ask what can judge an answer, and how much you trust it.
- A program can check the final answer (arithmetic, unit tests, a format, a proof checker): GRPO with that check as the reward. In this lesson's toy, accuracy on eight addition prompts goes from 23% to 99% in 60 steps of 8 answers each; this recipe is how DeepSeek-R1-Zero learned to reason from correct final answers alone.
- Only people can judge, and you have their ratings: a learned reward model with PPO, a KL leash, and a held-out measure of the real goal that you watch more closely than the reward.
- You have pairs but no budget for sampling: DPO.
- You can write the answers: supervised fine-tuning, and stop there.
- Whatever you pick, the policy optimises the reward you wrote, not the goal you meant. Before training, ask what a literal-minded optimiser would do with your reward, and measure the goal separately.
What it costs
Sampling is the bill: every training example is a full generation, and GRPO multiplies it by the group size, which is why PPO reuses each batch for several passes and why training loops run a fast inference engine beside the trainer. A learned baseline costs a second model (PPO's value network); GRPO replaces it with the group's own mean, which is why it was introduced as a way to cut PPO's memory. Noise costs steps: in the lesson's two-arm toy the gradient estimate's variance is 30.25 without a baseline and 0 with one, and with rewards offset by five points, 60% of training runs without a baseline lock onto the wrong arm against 0% with one. Reusing a batch too hard costs calibration: 50 passes over 16 pulls push one arm to 0.84 unclipped, 0.41 with PPO's clip at 0.2. And a KL leash costs a little reward on every answer (1.0 becomes 0.931 for an answer whose probability doubled, at β = 0.1) to buy fluency and safety from drift.
What breaks
- Reward hacking. The policy finds where the reward and the goal disagree. In the toy, a reward model fitted on answers of 1 to 4 sentences rates a 10-sentence answer 2.19, the worst answer of all; true quality rises from 0.80 to 0.93 and then falls to 0.75, below where it started, while the reward keeps climbing. Gao, Schulman and Hilton (2022) measured the same rise and fall at scale. "The reward went up" proves nothing; keep a held-out measure of the goal.
- Length bias and sycophancy. Raters prefer long, confident, flattering answers, so the reward model does too, so the model becomes that. Penalise length directly and rate the policy's current outputs, not stale ones.
- Gaming the checker. A coding model rewarded for passing tests learns to edit or special-case the tests. A verifier has no learned blind spot, but a buggy or bypassable one is a reward model with extra steps.
- Groups with nothing to teach. When every answer in a GRPO group is right or all are wrong, every advantage is zero: by the end of the toy run, 84% of groups are unanimous. Filter for prompts at the edge of the model's ability.
- Too loose a leash. As β falls, the best policy piles onto whatever the reward likes (99.995% on one action at β = 0.1 in the toy) and true quality collapses; too tight and nothing moves. Sweep it.
- Collapsed exploration. An arm whose probability hits zero is never tried again, so the policy can lock onto a mediocre answer early.
In the wild
InstructGPT (Ouyang et al., 2022) set the pattern of PPO against a learned reward model with a per-token KL penalty, and every chat assistant since inherits its habits. DeepSeekMath (Shao et al., 2024) introduced GRPO, reaching 51.7% on the MATH benchmark with a 7-billion-parameter model; DeepSeek-R1 (2025) trained reasoning with RL on rule-based rewards and no human-written reasoning traces. Hugging Face TRL's GRPOTrainer takes reward functions as plain Python callables or a reward model, samples 8 generations per prompt by default, and can generate with vLLM; OpenAI's model optimization guide lists reinforcement fine-tuning, where you supply the grader. Sutton and Barto's textbook and OpenAI's Spinning Up are the standard longer reads.
Go deeper
Level 2 builds it all on a three-armed slot machine: REINFORCE as one line of arithmetic, why a baseline removes noise without bias, PPO's ratio and clip on a table of four cases, the KL leash as a fee per answer, GRPO's advantages from a group of four, and a reward model that loves length, trained against until quality falls, each with a figure you can rerun. If you only needed to choose a training signal, you are done.
Level 2
How it works, from scratch
Think of teaching a dog to sit. You can't show it the right answer: you can only wait for it to try something and give it a treat when the something was good. Over many tries the dog does more of what earned treats and less of what didn't. Nobody ever told it what "sit" means; it worked it out from a score.
That is reinforcement learning (RL). An agent (the dog, or a language model) takes an action (sits, or writes an answer), the world hands back a reward (a treat, or a score from a grader), and the agent adjusts itself so that rewarded actions become more likely. The agent's current habits, written as a probability for every action, are called its policy.
Compare this with ordinary supervised learning (primer.ml.losses), where
every example comes with the correct answer attached. In RL there is no
answer key, only a score after the fact. That is exactly the situation a
language model is in once pretraining is over: for "prove this theorem" or
"write a helpful reply" there is no single correct text to copy, but a
checker or a judge can say how good an attempt was. RL is how the model
learns from those judgements (primer.ml.training_stages places it in the
training pipeline).
Figure 1 · Diagram
flowchart LR P[Policy<br/>a probability for every action] -->|sample| A[Action] A --> E[Environment<br/>a slot machine, a grader, a user] E --> R[Reward<br/>one number] R -->|nudge the policy| P
Chapter 1
A tiny worked example: three slot machines
The simplest RL problem is a row of slot machines, called a multi-armed bandit (a slot machine is a "one-armed bandit"). Machine A pays out 20% of the time, B 50% and C 80%, but the agent isn't told that. Each pull pays 1 or 0. The agent must find C by pulling and seeing what happens.
A policy here is three probabilities, one per arm. A policy that picks each arm a third of the time earns, on average, (0.2 + 0.5 + 0.8) / 3 = 0.5 per pull. A policy that always picks C earns 0.8. Learning means moving from the first policy to the second using nothing but the 1s and 0s.
This one number, the average reward a policy expects, is what RL maximises:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| one action: which arm to pull | A, B or C | |
| the policy's adjustable numbers (here, one score per arm, called logits) | (0, 0, 0) | |
| the policy: the probability of picking action , given . Here softmax of the logits | 1/3 each | |
| the average reward action pays (unknown to the agent) | 0.2, 0.5, 0.8 | |
| add up over every action | three terms | |
| the expected reward: what the policy earns per pull, on average | 0.5 |
In words: "the expected reward is each action's probability times its average payout, added up over the actions."
With the numbers: J = ⅓·0.2 + ⅓·0.5 + ⅓·0.8 = 0.5 for the uniform policy, and 1·0.8 = 0.8 for the policy that always pulls C.
Level 3: in Python
policy = [1/3, 1/3, 1/3]
# average payout of arms A, B, C
R = [0.2, 0.5, 0.8]
# J = Σ_a π(a) R(a)
round(sum(p * r for p, r in zip(policy, R)), 3) # → 0.5
round(sum(p * r for p, r in zip([0, 0, 1], R)), 3) # → 0.8
The agent can't compute J, because it doesn't know R. It can only sample: pull an arm, see a 1 or a 0. Every method below turns those samples into an estimate of which way to move θ to make J bigger. Trying an arm you're unsure of is called exploration; sticking with the best arm so far is exploitation. A sampled policy does some of both automatically, as long as no arm's probability has collapsed to zero.
In code: Bandit hides the win chances and pays out one pull at a time; expected_reward is the formula above.
Chapter 2
Policy gradients: do more of what worked (REINFORCE)
Everyday picture A football coach reviews the tape after a match. For every play that led to a goal, they tell the team "a bit more of that"; for plays that went nowhere, nothing. They don't need to know why the play worked. Repeat over hundreds of matches and the team drifts towards the plays that score.
Tiny worked example Start with logits (0, 0, 0), so each arm has probability ⅓. The agent pulls C and wins: reward 1. The rule, explained next, says: add to each logit learning rate × reward × (1 if it's the chosen arm, else 0, minus that arm's probability).
| Arm | Chosen? | 1[chosen] − π | × reward 1 × rate 0.5 | New logit | New probability |
|---|---|---|---|---|---|
| A | no | 0 − ⅓ = −0.333 | −0.167 | −0.167 | 0.274 |
| B | no | 0 − ⅓ = −0.333 | −0.167 | −0.167 | 0.274 |
| C | yes | 1 − ⅓ = +0.667 | +0.333 | +0.333 | 0.452 |
One lucky pull moved C from 33% to 45%. Had the pull paid 0, nothing would have moved. Had the agent pulled A and won (A wins sometimes too), A would have gone up instead. The rule is noisy, one pull at a time, but on average the arm that wins most gets pushed up most.
Figure 3 · Diagram
flowchart LR L[Logits θ] --> S[softmax<br/>probabilities π] S -->|sample| A[Action a] A --> ENV[Pull the arm] --> R[Reward R] S --> G["∇ log π(a)<br/>= one-hot(a) − π"] A --> G G --> M["× R × learning rate"] R --> M M -->|add to| L
∇ log π box), then scale that direction
by how good the outcome was. A big reward is a big step towards repeating the
action; zero reward is no step.The log-probability trick, decoded
We want the gradient of J: for each logit, how much J rises if the
logit rises a little (see primer.notation for gradients from scratch).
The difficulty is that J is an average over actions we can only sample. The
trick rewrites the gradient as an average too, so a sample estimates it:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| "gradient with respect to θ": one slope per logit, collected into a vector | 3 slopes | |
| expected value: the average of the bracket when is sampled from the policy | average over pulls | |
| "drawn from" | ||
| the natural logarithm, the undo button for | ||
| the direction in logit space that makes action more likely, fastest | for C | |
| the -th logit (θ is the list of logits) | ||
| "partial derivative": the slope along one logit, holding the others still | ||
| 1 if is the chosen action, else 0 | (0, 0, 1) | |
| the probability of action | ⅓ |
In words: "the direction that raises expected reward is, on average, the direction that makes the sampled action more likely, weighted by the reward it earned. For a softmax policy, that direction is 'one for the chosen action, minus every action's probability'."
Why is this true? Because the slope of a probability equals the probability
times the slope of its log (, the chain
rule applied to log). So $\nabla J = \sum_a \nabla\pi(a) R(a) = \sum_a \pi(a)
\nabla\log\pi(a) R(a)\pi(a)$ is an average over
samples from π. The update is then plain gradient ascent:
,
with learning rate (primer.ml.optimizers).
With the numbers: at the uniform policy the true gradient is (−0.1, 0, +0.1): push C up, A down, leave B (which pays exactly the average) alone. One sample, "pulled C, got 1", estimates it as 1 × (−⅓, −⅓, ⅔). The step with α = 0.5 gives logits (−0.167, −0.167, 0.333) and probabilities (0.274, 0.274, 0.452), as in the table.
In Python:
import math
z = [0.0, 0.0, 0.0]
pi = [math.exp(z_k) / sum(math.exp(v) for v in z) for z_k in z]
# the true gradient: Σ_a π(a) R(a) (1[k=a] − π(k)), for each logit k
R = [0.2, 0.5, 0.8]
true_grad = [sum(pi[a] * R[a] * ((k == a) - pi[k]) for a in range(3)) for k in range(3)]
[round(g, 3) for g in true_grad] # → [-0.1, 0.0, 0.1]
# one sample: pulled C (a = 2), reward 1
a, reward, alpha = 2, 1.0, 0.5
grad_log_pi = [(k == a) - pi[k] for k in range(3)]
[round(g, 3) for g in grad_log_pi] # → [-0.333, -0.333, 0.667]
z = [z_k + alpha * reward * g for z_k, g in zip(z, grad_log_pi)]
[round(z_k, 3) for z_k in z] # → [-0.167, -0.167, 0.333]
[round(math.exp(z_k) / sum(math.exp(v) for v in z), 3) for z_k in z] # → [0.274, 0.274, 0.452]
This algorithm is called REINFORCE (Williams, 1992). Run it for 500 pulls and the policy finds arm C:
Figure 2 · Drawn from the lesson's code
REINFORCE on the three-armed bandit: the probability of arm C climbs from a third to about 0.96 within 500 pulls, and the average reward rises from 0.5 towards 0.8
Why it matters in practice. A language model is exactly this kind of policy, with a vocabulary of tokens as its arms, and one sampled answer is a string of sampled tokens. REINFORCE applies unchanged: sum the log probabilities of every token in the answer, and scale the gradient by the answer's reward. Every method below (PPO, GRPO) is REINFORCE with repairs.
In code: grad_log_prob is one-hot minus the probabilities, reinforce_step is one update, worked_reinforce_step is the table above, and train_reinforce runs the whole loop against a Bandit.
Chapter 3
Variance and baselines: grade on a curve
Everyday picture A teacher whose class all scores between 90 and 100 learns nothing by being told "you got a 92". What matters is whether 92 is above or below the class average. Raw scores that are all large and positive make every attempt look good; only the difference from typical tells you which way to go.
Tiny worked example Two arms, a 50/50 policy, and every pull pays a lot: arm 1 always pays 10, arm 2 always pays 12. Arm 2 is better, so the logit of arm 2 should rise. Look at the REINFORCE estimate for that logit:
| Pulled | Reward | 1[arm 2] − π(arm 2) | Estimate (no baseline) | Estimate (baseline 11) |
|---|---|---|---|---|
| arm 1 | 10 | 0 − 0.5 = −0.5 | 10 × −0.5 = −5 | (10 − 11) × −0.5 = +0.5 |
| arm 2 | 12 | 1 − 0.5 = +0.5 | 12 × +0.5 = +6 | (12 − 11) × +0.5 = +0.5 |
Without a baseline the estimate is −5 or +6 depending on the coin flip. It averages to +0.5, the right answer, but any single sample points the wrong way half the time, and violently. Subtract the average reward, 11, first and every sample says +0.5. Same average, no noise at all.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the baseline: any number that doesn't depend on which action was taken; usually the average reward | 11 | |
| the advantage: how much better action did than typical | −1 for arm 1, +1 for arm 2 | |
| everything else | as in the REINFORCE formula above |
In words: "scale each step by how much better than typical the action did, not by its raw reward."
Why is it allowed? Because the baseline's contribution averages to zero: $\mathbb{E}[, b, \nabla \log \pi(a)] = b \sum_a \nabla \pi(a) = b, \nabla \sum_a \pi(a) = b, \nabla 1 = 0$. Probabilities always add to 1, so pushing all of them up is impossible; the baseline only removes noise, never signal.
With the numbers: without a baseline the estimate's variance (the
average squared distance from its mean, see primer.notation) is
(25 + 36)/2 − 0.5² = 30.25. With b = 11 it is 0.
In Python:
# arm 1 pays 10, arm 2 pays 12; the 1[arm 2] − π(arm 2) factor for each pull
pulls = [(10, -0.5), (12, +0.5)]
def mean_and_variance(b):
estimates = [(reward - b) * direction for reward, direction in pulls]
mean = sum(estimates) / 2
return mean, sum((e - mean) ** 2 for e in estimates) / 2
mean_and_variance(b=0) # → (0.5, 30.25)
mean_and_variance(b=11) # → (0.5, 0.0)
Figure 5 · Diagram
flowchart LR
R[Reward R] --> MINUS["R − b"]
B["Baseline b<br/>average reward so far"] --> MINUS
MINUS --> ADV{Advantage A}
ADV -->|positive: better than typical| UP[make the action<br/>more likely]
ADV -->|negative: worse than typical| DOWN[make the action<br/>less likely]
To see it matter, give every arm of our bandit 5 extra points: rewards are now 5 or 6 instead of 0 or 1, and nothing about which arm is best has changed.
Figure 4 · Drawn from the lesson's code
With rewards offset by 5, the gradient estimate's variance falls from about 20 to 0.15 with a baseline, and 20 training runs all find arm C with it, while without it most runs lock onto the wrong arm
Why it matters in practice. Every practical policy-gradient method uses a baseline. PPO learns one with a second network, the value network (or critic), which predicts the expected reward from each state. GRPO, below, gets one for free by comparing several answers to the same prompt.
In code: gradient_estimate_stats computes the exact mean and variance of the estimate, sampled_gradient_variance measures it from real pulls, and train_reinforce subtracts a running average when baseline=True.
Chapter 4
PPO: take several steps, but never too far
Everyday picture A chef tests a new recipe on one evening's diners. It would be wasteful to use their comments for just one small tweak, so the chef makes several rounds of changes from the same comment cards. But the further the recipe drifts from what the diners actually ate, the less their comments apply, so the chef caps each change: never more than 20% more or less of any ingredient per round.
For a language model, sampling answers is the expensive part (every answer is a full generation), so PPO (Proximal Policy Optimization) reuses each batch of answers for several gradient steps. It needs a way to tell how far the policy has moved since the batch was sampled, and a brake.
Tiny worked example The probability ratio compares the policy now with the policy that generated the sample. A ratio of 1.5 means the current policy is 50% more likely to produce that answer than when it was sampled. With a clip range ε = 0.2 the ratio is allowed to count only between 0.8 and 1.2:
| Ratio ρ | Advantage A | ρ·A | clip(ρ, 0.8, 1.2)·A | min of the two | What happened |
|---|---|---|---|---|---|
| 1.1 | +3 | 3.3 | 3.3 | 3.3 | inside the band: plain REINFORCE |
| 1.5 | +2 | 3.0 | 1.2 × 2 = 2.4 | 2.4 | good action already boosted enough: gain capped |
| 0.5 | −1 | −0.5 | 0.8 × −1 = −0.8 | −0.8 | bad action already cut enough: capped |
| 1.5 | −1 | −1.5 | 1.2 × −1 = −1.2 | −1.5 | bad action made more likely: full penalty |
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| one sample in the batch; for a language model, one token of one answer | row of the table | |
| the state: what the policy saw before acting (the prompt plus the tokens so far) | a prompt | |
| the action taken (the token generated) | an answer | |
| the policy as it was when the batch was sampled, frozen | ||
| the policy now, after some steps on this batch | ||
| the probability ratio, new over old | 1.5 | |
| the advantage of that sample | +2 | |
| the clip range, typically 0.1 to 0.3 | 0.2 | |
| , but pushed back to or if it falls outside | clip(1.5, 0.8, 1.2) = 1.2 | |
| the smaller of the two | min(3.0, 2.4) = 2.4 | |
| the average over all samples in the batch | ||
| the objective PPO climbs |
In words: "for each sample, take the ratio-weighted advantage, but once the ratio has moved more than ε from 1 in the direction the advantage wants, stop counting further movement; and always take the more pessimistic of the clipped and unclipped versions."
The slope of ρ·A is exactly REINFORCE's gradient scaled by ρ (because ), so inside the band PPO is REINFORCE with importance weighting. Outside it, the clipped term is flat: that sample stops pushing.
With the numbers: the four rows of the table, computed:
In Python:
def clip(x, lo, hi):
return max(lo, min(x, hi))
def L_clip(rho, A, eps=0.2):
return min(rho * A, clip(rho, 1 - eps, 1 + eps) * A)
[round(L_clip(rho, A), 2) for rho, A in [(1.1, 3), (1.5, 2), (0.5, -1), (1.5, -1)]] # → [3.3, 2.4, -0.8, -1.5]
# past 1 + ε with a positive advantage, a higher ratio earns nothing more
L_clip(1.3, 2) == L_clip(1.6, 2) # → True
Figure 7 · Diagram
flowchart TB
OLD[Policy π_old] -->|generate a batch| B[Answers + rewards]
B --> V[Value network<br/>predicts expected reward]
V --> ADV[Advantages A_t]
B --> ADV
ADV --> LOOP
subgraph LOOP["Several passes over the same batch"]
RATIO["ratio ρ = π_θ / π_old"] --> CLIP["clip to 1 ± ε, take the min"]
CLIP --> KL["subtract β × KL to the reference model"]
KL --> STEP[gradient step on θ]
STEP --> RATIO
end
LOOP -->|π_θ becomes the new π_old| OLD
To see the clip work, take one batch of 16 pulls from the bandit (C won all four of its pulls, A lost all four) and make 50 passes over it:
Figure 6 · Drawn from the lesson's code
Two panels: the clipped objective is flat outside the 0.8 to 1.2 band on the side the advantage favours; over 50 passes the unclipped ratio climbs past 2.5 while the clipped one levels off near 1.24
In code: ppo_clipped_objective is the formula, clipped_policy_update makes several passes over one batch with or without the clip, and ppo_drift_experiment is the 50-pass comparison.
The KL leash: stay close to where you started
Everyday picture A dog on a long leash can explore, but it can't run off a cliff. In RL for language models, the leash ties the policy to a frozen copy of the model it started from, the reference model (usually the model after supervised fine-tuning).
Tiny worked example An answer earns reward 1.0. The policy now gives it probability 0.6; the reference gave it 0.3. With leash strength β = 0.1, the reward actually used for training is 1.0 − 0.1 × ln(0.6 / 0.3) = 1.0 − 0.1 × 0.693 = 0.931. The policy pays a small fee for having doubled that answer's probability.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the reward from the grader or reward model | 1.0 | |
| the reward after the leash's fee | 0.931 | |
| leash strength: how much a unit of drift costs | 0.1 | |
| the frozen reference model's probability for the answer | 0.3 | |
| how much more (positive) or less (negative) likely the policy makes this answer than the reference | ln 2 = 0.693 | |
KL divergence: the average of that log ratio over the policy's own answers; zero only when the two agree (decoded in primer.ml.training_stages) |
In words: "each answer's reward is docked in proportion to how much more likely the policy has made it than the reference did; on average, that fee is β times the KL divergence between the two."
With the numbers: 1.0 − 0.1 × ln 2 = 0.931. Had the policy halved the answer's probability instead (0.15), the log ratio would be −0.693 and the reward would rise to 1.069: the leash pulls both ways.
Level 3: in Python
import math
R, beta = 1.0, 0.1
pi, pi_ref = 0.6, 0.3
# R' = R − β log(π / π_ref)
round(R - beta * math.log(pi / pi_ref), 3) # → 0.931
round(R - beta * math.log(0.15 / pi_ref), 3) # → 1.069
Why it matters in practice. This is the "penalty for drifting" in the
RLHF loop of primer.ml.training_stages, and the same β appears in DPO. It
stops the policy from forgetting fluent language while it chases reward, and
it is the first line of defence against reward hacking, below.
In code: kl_penalised_reward is R′; clipped_policy_update adds the leash's gradient when given a reference policy and a β, and primer.ml.training_stages.kl_divergence computes the KL itself.
Chapter 5
GRPO: compare answers to the same question
Everyday picture Instead of hiring an examiner to predict how hard each exam question is, a teacher gives the same question to eight students and marks each answer relative to the others on that question. On an easy question, getting it right is expected and earns little credit; on a hard one, the only right answer stands out.
PPO's baseline comes from the value network, a second model as large as the policy that has to be trained alongside it. GRPO (Group Relative Policy Optimization) throws the value network away. For each prompt it samples a group of answers and uses the group's own average as the baseline.
Tiny worked example The prompt is "3 + 4 =". The model samples four answers and a checker scores them 1 if the answer is 7, else 0.
| Group rewards | Mean | Std | Advantages |
|---|---|---|---|
| 1, 0, 0, 1 | 0.5 | 0.5 | +1, −1, −1, +1 |
| 1, 0, 0, 0 | 0.25 | 0.433 | +1.73, −0.58, −0.58, −0.58 |
| 1, 1, 1, 1 | 1 | 0 | 0, 0, 0, 0 |
A lone right answer in a mostly wrong group earns a big advantage: it's rare, so it's strong evidence. A group that is all right (or all wrong) earns nothing: there is no contrast, so there is nothing to learn from.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the group size: answers sampled per prompt | 4 | |
| which answer in the group | 1 … 4 | |
| the reward for answer | 1, 0, 0, 0 | |
| the group's average reward: the baseline | 0.25 | |
| the group's standard deviation, the typical distance from the mean (square root of the variance) | 0.433 | |
| answer 's advantage, shared by every token of that answer | +1.73 |
In words: "an answer's advantage is how far its reward sits above the group's average, measured in units of the group's spread."
With the numbers: for (1, 0, 0, 0): mean 0.25, variance (0.75² + 3 × 0.25²) / 4 = 0.1875, std √0.1875 = 0.433, so the right answer gets 0.75 / 0.433 = 1.73 and each wrong one −0.25 / 0.433 = −0.58.
Level 3: in Python
import statistics
def advantages(r):
mu, sd = statistics.mean(r), statistics.pstdev(r)
return [round((r_i - mu) / sd, 2) if sd else 0.0 for r_i in r]
advantages([1, 0, 0, 1]) # → [1.0, -1.0, -1.0, 1.0]
advantages([1, 0, 0, 0]) # → [1.73, -0.58, -0.58, -0.58]
advantages([1, 1, 1, 1]) # → [0.0, 0.0, 0.0, 0.0]
The rest of GRPO is PPO: the same ratio, the same clip, the same KL leash to a reference model, averaged over the group. (Some implementations divide by the sample standard deviation, with G − 1, rather than the population one; the idea is identical.)
Figure 9 · Diagram
flowchart LR
subgraph PPO["PPO"]
P1[Prompt] --> A1[one answer]
A1 --> RM1[reward]
A1 --> VN[value network<br/>a second big model]
RM1 --> AD1[advantage = reward − value]
VN --> AD1
end
subgraph GRPO["GRPO"]
P2[Prompt] --> G1[answer 1] & G2[answer 2] & G3[answer ...] & G4[answer G]
G1 & G2 & G3 & G4 --> VER[verifier<br/>checks each answer]
VER --> NORM[normalise within the group<br/>mean and std]
NORM --> AD2[advantage per answer]
end
Verifiable rewards, and why GRPO trains reasoning
A verifiable reward comes from a check that can't be argued with: does the arithmetic equal 7, do the unit tests pass, does the proof check. No learned reward model, so no learned blind spots (see reward hacking, next). Here is GRPO on a toy "language model" that answers eight addition prompts with a single digit token. It starts out about 23% accurate and leans towards off-by-one mistakes.
Figure 8 · Drawn from the lesson's code
GRPO on eight addition prompts: accuracy climbs from 23% to 99% in 60 steps, while the share of prompts whose whole group agreed, and so taught nothing, rises from a few percent to over 80%
Why it matters in practice. This recipe, a verifier on the final
answer plus GRPO, is how DeepSeek-R1-Zero learned to reason: rewarded only
for correct final answers (and a required format), the model learned by
itself to write longer chains of thought, to check its work and to back
up from mistakes, because those behaviours raised the chance of a correct
final answer. Each token in a long chain of thought shares its answer's
advantage, so the whole chain is reinforced or discouraged together. See
primer.ml.reasoning for what that training produces.
In code: group_advantages is the formula, verify is the checker, pretrained_logits is the weak starting model, and train_grpo runs the loop.
Chapter 6
Reward hacking: the score is not the goal
Everyday picture A school pays tutors by the number of pages of homework feedback they write. Feedback gets longer, not better. An RL agent is the most literal-minded employee imaginable: it optimises the number you wrote down, not the thing you meant. When those two differ, it finds the difference. This is reward hacking (also called specification gaming), and it is Goodhart's law in code: when a measure becomes a target, it ceases to be a good measure.
Tiny worked example The goal is a good answer, and good answers here are about 4 sentences long. People rated answers of 1 to 4 sentences, and on those, longer really was better. A reward model fitted to their ratings learns a straight line: "each sentence is worth 0.19 points". Nobody ever rated a 10-sentence answer, so nobody told the reward model it was bad:
| Sentences | 1 | 2 | 3 | 4 | 6 | 8 | 10 |
|---|---|---|---|---|---|---|---|
| True quality | 0.44 | 0.75 | 0.94 | 1.00 | 0.75 | 0.00 | −1.25 |
| Reward model | 0.50 | 0.69 | 0.88 | 1.06 | 1.44 | 1.81 | 2.19 |
The reward model's favourite answer, 10 sentences, is the worst one.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the answer's length in sentences | 1 … 10 | |
| the reward model's score for a length- answer (the hat marks an estimate) | 2.19 at n = 10 | |
| the rated examples: a length and its true quality | (1, 0.44), …, (4, 1.0) | |
| their averages (the bar means "mean") | 2.5, 0.781 | |
| the fitted slope: points per extra sentence | 0.1875 | |
| the fitted intercept | 0.3125 |
In words: "the reward model is the straight line that best fits the ratings it saw: its slope is how length and quality moved together in the data, and it passes through the average point."
With the numbers: the deviations of length are (−1.5, −0.5, 0.5, 1.5) and of quality (−0.344, −0.031, 0.156, 0.219). Their products add to 0.9375 and the squared length deviations to 5, so w₁ = 0.1875 and w₀ = 0.781 − 0.1875 × 2.5 = 0.3125. At n = 10 it scores 2.19.
Level 3: in Python
n = [1, 2, 3, 4]
q = [1 - ((n_i - 4) / 4) ** 2 for n_i in n]
q # → [0.4375, 0.75, 0.9375, 1.0]
n_bar, q_bar = sum(n) / 4, sum(q) / 4
w1 = sum((a - n_bar) * (b - q_bar) for a, b in zip(n, q)) / sum((a - n_bar) ** 2 for a in n)
w0 = q_bar - w1 * n_bar
(w0, w1) # → (0.3125, 0.1875)
# the reward model's score for a 10-sentence answer, and its true quality
(w0 + w1 * 10, 1 - ((10 - 4) / 4) ** 2) # → (2.1875, -1.25)
Figure 10 · Drawn from the lesson's code
True quality peaks at 4 sentences and falls below zero past 8, while the reward model's straight line, fitted on lengths 1 to 4, keeps rising to 2.19 at 10 sentences
Figure 12 · Diagram
flowchart LR GOAL[What we want<br/>helpful answers] -->|people rate a sample| DATA[Ratings] DATA -->|fit| RM[Reward model<br/>a proxy for the goal] RM -->|reward| OPT[RL optimiser] OPT --> POL[Policy] POL -->|drifts to where<br/>proxy and goal disagree| GAP[Gap:<br/>high reward, low quality] KL[KL leash] -.->|limits drift| POL VER[Verifiable reward] -.->|no learned gap| OPT
Now train the starting model (which writes 2 or 3 sentences, a little too short) against each reward:
Figure 11 · Drawn from the lesson's code
Over 300 steps, optimising the flawed reward first raises true quality to 0.93 then drives it down to 0.75, below where it started, while a KL leash holds it near 0.84 and the verifiable reward climbs to 1.0; the right panel shows the best leashed policy's true quality peaking at a moderate KL and collapsing as the leash loosens
The right-hand panel uses a closed form: the best policy under a KL leash has an exact formula.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the policy that maximises | (0.731, 0.269) | |
| the reference model's probability for action | (0.5, 0.5) | |
| the reward for action | (1, 0) | |
| leash strength | 1 | |
| a boost that grows with reward; a small β makes it enormous | ||
| a counter over every action, so adds up all of them | ||
| the total, so the probabilities add to 1 | 1.859 |
In words: "the best leashed policy starts from the reference and multiplies each action's probability by e to the power reward over β, then rescales so everything adds to 1."
With the numbers: two actions, reference (0.5, 0.5), rewards (1, 0), β = 1: weights 0.5 × 2.718 = 1.359 and 0.5 × 1 = 0.5, total 1.859, so π* = (0.731, 0.269). At β = 0.1 the boost is e¹⁰ ≈ 22,026 and π* puts 99.995% on the rewarded action; as β grows, π* returns to the reference.
Level 3: in Python
import math
ref, R = [0.5, 0.5], [1.0, 0.0]
def best_policy(beta):
w = [p * math.exp(r / beta) for p, r in zip(ref, R)]
return [round(w_a / sum(w), 5) for w_a in w]
best_policy(1.0) # → [0.73106, 0.26894]
best_policy(0.1) # → [0.99995, 5e-05]
best_policy(100.0) # → [0.5025, 0.4975]
This is the same formula DPO starts from (primer.ml.training_stages): it
is why β means the same thing in RLHF and DPO.
Why it matters in practice. Reward hacking shows up wherever RL does. A boat-racing game agent that learned to circle forever collecting bonus targets instead of finishing the race. RLHF'd chat models that learned long, confident, flattering answers score well with raters (length bias and sycophancy). Coding models rewarded for passing tests that learned to edit or special-case the tests. The defences, strongest first:
- Verifiable rewards where the task allows: run the tests, check the answer. A check has no learned blind spot (though a buggy check does).
- Better reward models: rate the policy's current outputs and retrain, so the reward model sees the regions the policy is exploring; use ensembles; penalise known exploits such as length directly.
- A KL leash to keep the policy near the data the reward was fit on.
- Watch a held-out measure of the real goal (human review, a separate
evaluation set, see
primer.agents.evals) and stop when it turns down, even if the reward is still climbing.
In code: true_quality is the goal, fit_reward_model and proxy_reward are the flawed reward, optimise_lengths trains against either with an optional leash, and kl_regularised_optimum with leash_sweep gives the closed-form best policy for each β.
Test yourself
9 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1How does reinforcement learning differ from supervised learning?Think it through, then reveal
Supervised learning is given the correct output for every input and learns to copy it. Reinforcement learning is given only a score for the output it produced, so it must try things, see how they score and shift probability towards what scored well. It fits tasks where judging an answer is easy but writing the perfect one is not.
Question 2What is the log-probability trick, and why is it needed?Think it through, then reveal
The gradient of expected reward, Σ ∇π(a) R(a), can't be computed without knowing every action's reward. Rewriting ∇π = π ∇log π turns it into an average over actions sampled from the policy, E[R ∇log π(a)], so each sampled action and its reward give an unbiased estimate of the gradient.
Question 3Why does subtracting a baseline not change the expected gradient?Think it through, then reveal
Because E[b ∇log π(a)] = b ∇ Σ π(a) = b ∇ 1 = 0: probabilities always add to 1, so the baseline's push averages to nothing. It only removes the noise that comes from rewards being large or all the same sign.
Question 4In PPO, what is the probability ratio, and what does clipping it do?Think it through, then reveal
The ratio is the current policy's probability of a sampled action divided by the probability under the policy that sampled it; it measures how far the policy has moved on that sample. Clipping stops counting movement beyond 1 ± ε in the direction the advantage favours, so reusing a batch for several steps can't push the policy far from where the data came from.
Question 5Why does PPO's objective take the minimum of the clipped and unclipped terms?Think it through, then reveal
To stay pessimistic. Gains are capped once the ratio leaves the band, but if a step made a bad action more likely, the full penalty still applies, so the objective never rewards a harmful move.
Question 6How does GRPO get a baseline without a value network?Think it through, then reveal
It samples several answers to the same prompt and uses their mean reward as the baseline, dividing by their standard deviation to set the scale. Each answer is judged against its siblings, which saves training and serving a second model the size of the policy.
Question 7What happens in GRPO when every answer in a group gets the same reward?Think it through, then reveal
Every advantage is zero, so that prompt contributes no gradient. Prompts that are always solved or never solved teach nothing; learning comes from prompts at the edge of the model's ability.
Question 8What is reward hacking, and why does a KL penalty help against it?Think it through, then reveal
Reward hacking is the policy maximising the reward as written while the real goal gets worse, usually by finding inputs where a learned reward model is wrong. The KL penalty charges the policy for drifting from the reference model, which keeps it near the kind of outputs the reward model was trained on, where the reward is still trustworthy.
Question 9Why are verifiable rewards attractive for training reasoning?Think it through, then reveal
A program that checks the final answer (or runs the tests) has no learned blind spots to exploit and costs nothing to label, so RL can run for a long time against it without the reward drifting away from correctness.
Primary sources
The papers behind this lesson
Introduced REINFORCE, the log-probability policy gradient with a baseline.
Read the annotated companion →The paper ↗Introduced the clipped probability-ratio objective that lets each batch be reused for several safe steps.
Read the annotated companion →The paper ↗Used PPO with a per-token KL penalty to a reference model to tune a language model against a learned reward model.
Read the annotated companion →The paper ↗Introduced GRPO, replacing PPO's value network with group-relative advantages.
Read the annotated companion →The paper ↗Showed GRPO with rule-based, verifiable rewards alone can teach a model to produce long, self-checking chains of thought.
Read the annotated companion →The paper ↗Measured how true quality rises and then falls as a policy is optimised further against a learned reward model.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Sutton & Barto, Reinforcement Learning: An Introduction (2nd edition, free online): http://incompleteideas.net/book/the-book-2nd.html
- OpenAI, Spinning Up in Deep RL: https://spinningup.openai.com/
- Andrej Karpathy, Deep Reinforcement Learning: Pong from Pixels: http://karpathy.github.io/2016/05/31/rl/
- Lilian Weng, Policy Gradient Algorithms: https://lilianweng.github.io/posts/2018-04-08-policy-gradient/
- Hugging Face TRL, GRPO Trainer: https://huggingface.co/docs/trl/grpo_trainer
- Amodei et al., Concrete Problems in AI Safety (2016), section on reward hacking: https://arxiv.org/abs/1606.06565
- Schulman et al., Proximal Policy Optimization Algorithms (2017): https://arxiv.org/abs/1707.06347
- Shao et al., DeepSeekMath (2024), which introduces GRPO: https://arxiv.org/abs/2402.03300
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.