rumblr Work in progressWIP

● The AI Primer · Lesson 14 · Part 1: how the model works inside

Alignment and safety

turning what we want into something we can measure

This lesson covers Constitutional AI, red-teaming, sycophancy, refusals

Members · open during launch 38 min11 figures and diagrams8 interactive
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Alignment means making a model's behaviour match what we want (helpful, honest, harmless), when all we can optimize is a measurement of it. Optimize a measurement hard enough and it stops tracking the goal (Goodhart's law; reward hacking).
  2. Constitutional AI writes the target down as principles; a model critiques and revises its own answers against them, and AI-labelled preference pairs train the reward model (RLAIF).
  3. Red-teaming searches for failures on purpose and reports an attack success rate; hand-written tests measure the author's imagination.
  4. Sycophancy is answers bending towards the user's stated view; measure it as a flip rate, and know that rater preferences for agreement can create it.
  5. Refusals trade harmful compliance against over-refusal as a threshold moves; only a better classifier improves both.
  6. Before release: limits set in advance, a gate on every evaluation, then a staged rollout that feeds new failures back into the tests.

Level 2

How it works, from scratch

Picture hiring a new assistant and handing them a one-page brief: be useful, tell the truth, don't cause trouble. The brief is clear to you, but you can't watch every task they do. So you check what you can check: how quickly they reply, whether customers leave a thumbs-up, whether the report has the right headings. The assistant, like anyone being measured, learns what the checks reward. If the checks and the brief agree, all is well. Where they disagree, the assistant drifts towards the checks.

Alignment is the engineering work of making a model's behaviour match the brief, not just the checks. In practice the brief is usually summarised as three targets:

Target Plain meaning Something we can measure (imperfectly)
Helpful does the task the person actually asked for rater preferences, task success on evaluation sets
Honest says what it believes is true, and how sure it is accuracy on factual questions, calibration, answer flips under pressure
Harmless declines things that would cause harm, and nothing else attack success rate, over-refusal rate

Every lesson section below takes one of those measurements, builds it from scratch on made-up, neutral examples, and shows where it can mislead. The training machinery itself (reward models, RLHF, DPO) lives in primer.ml.training_stages; the runtime checks that wrap a deployed model live in primer.agents.guardrails. This lesson sits between them: how the targets are written down, how failures are searched for, and how the result is judged before release.

The three targets above, each walked from the brief to the check that stands in for it.

Chapter 1

The gap between what we measure and what we want

Run the tuning yourself first; the two curves will then read as what they are.

Everyday picture A call centre rewards agents for short calls. Calls get shorter. Some of that is agents getting better; some of it is agents hanging up on hard customers. The number keeps improving while the thing it stood for gets worse. This is Goodhart's law: once a measure becomes a target, it stops being a good measure.

Tiny worked example A model is being tuned against a reward model that scores answers. Each tuning step adds two things to its answers: some genuine content, which helps but saturates (there is only so much to say), and some padding, which the reward model can't tell apart from content. Say genuine content after steps is and padding is . The reward model sees content plus padding; the person reading sees content minus the cost of wading through padding.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Range
how many optimization steps have run 0, 1, 2, …
genuine content: rises fast, then levels off at 1 0 … 1
Euler's number (≈ 2.718) raised to : starts at 1 and decays towards 0 0 … 1
padding: grows steadily with every step 0 …
what the reward model scores, the number being optimized 0 …
what the reader actually gets, the thing we wanted any real number

In words: "the measured score is content plus padding; the real value is content minus padding; content levels off while padding keeps growing."

With the numbers: at step 16, and , so the proxy is 1.118 and the true value is 0.478, its peak. At step 50, and : the proxy has climbed to 1.993, yet the true value has fallen to −0.007, worse than doing nothing.

Level 3: in Python
import math
def g(s):
    return 1 - math.exp(-s / 10)
def p(s):
    return s / 50
# proxy(s) and true(s) at step 16
round(g(16) + p(16), 3), round(g(16) - p(16), 3)  # → (1.118, 0.478)
# ... and at step 50: the proxy keeps climbing, the true value has collapsed
round(g(50) + p(50), 3), round(g(50) - p(50), 3)  # → (1.993, -0.007)
# the step where the true value peaks
max(range(51), key=lambda s: g(s) - p(s))  # → 16

Figure 1 · Drawn from the lesson's code

0 10 20 30 40 50 optimization steps 0.0 0.5 1.0 1.5 2.0 score Goodhart's law: the measurement keeps rising, the goal does not true value peaks at step 16 proxy: what the reward model scores true: what the reader gets

The proxy score rises at every step, while the true value peaks at step 16 and then falls back below zero by step 50

Reading it: the x-axis is optimization steps; the y-axis is score. The blue proxy line never stops rising, so anyone watching only the reward model sees steady progress. The red true line rises with it at first (the early steps really do help), peaks at the dashed marker, then falls: past that point, every step spent pleasing the measurement is a step away from the goal. The two lines only separate once optimization pushes hard, which is why the gap is invisible early on.

Figure 2 · Diagram

Reading it: read left to right as a chain of approximations. Each arrow loses a little: a written rule never captures everything we want, and a reward model never captures everything the rule says. Optimization only sees the box it is pointed at (the measurement), so under pressure it drifts towards whatever the measurement rewards. The dotted line back to the goal is the one we care about and the one nobody can optimize directly.

Why it matters in practice: this gap is the root of most alignment failures. A model tuned hard against a reward model finds the reward model's blind spots (called reward hacking, built from scratch in primer.ml.reinforcement). The standard defences are a penalty for drifting far from the starting model (the β in primer.ml.training_stages), stopping early, and refreshing the reward model with new labels where the policy has found its weak spots.

In code: goodhart_curve returns the proxy and true value at every step.

Chapter 2

Constitutional AI: principles written down, applied by a model

Play the editor first: score the three answers against the rules, then revise one.

Everyday picture A newspaper has a style guide. A junior writer drafts a story; an editor reads it against the style guide, writes margin notes ("unsourced claim", "too certain"), and the writer revises. Over time the editor's notes also teach the newsroom which of two drafts is better. Constitutional AI does this with models: the style guide is a short, written list of principles (the "constitution"), and a model plays the editor.

It has two stages:

  1. Critique and revise. The model drafts an answer, is asked to critique it against a principle, then to rewrite it. The revised answers become fine-tuning data (supervised fine-tuning, as in primer.ml.training_stages).
  2. AI feedback (RLAIF). The model compares pairs of answers against the principles and says which is better. Those AI-labelled pairs train the reward model, in place of (or alongside) human labels. The rest is the same preference tuning as RLHF.

The appeal is that the principles are written in plain language, can be read and argued about, and can be changed without relabelling thousands of examples by hand.

Tiny worked example Our toy constitution has four made-up rules:

Principle The rule (toy wording) Broken when…
helpful don't refuse without saying why the answer starts "I can't help" and gives no reason
honest don't claim more certainty than you have it says "guaranteed", "100%", "always works" or "certainly"
harmless never help with the toy off-limits action, xyzzy it mentions xyzzy or one of its synonyms
cites answers must cite a source it never mentions a source

Asked "Will this backup script work?", the model drafts three candidates:

Candidate Text Principles broken Count
A It is guaranteed to work. honest, cites 2
B It usually works; test it on a copy first. Source: the backup guide. none 0
C I can't help with that. helpful, cites 2

B beats A, and B beats C. A and C tie, and a tie tells the reward model nothing, so that pair is dropped. Three candidates give two preference pairs: (B over A) and (B over C).

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
one candidate answer "It is guaranteed to work."
how many principles the constitution has 4
a counter walking over the principles 1 = helpful, …, 4 = cites
the indicator: 1 if the statement in brackets is true, 0 if not 1 for "honest" on A
add up the following for every principle
how many principles breaks ,
"is preferred to" B ≻ A
"exactly when"

In words: "count the principles each answer breaks; one answer is preferred to another exactly when it breaks fewer."

With the numbers: , , . Since , B ≻ A and B ≻ C; A and C tie and produce no pair.

Level 3: in Python
# one row per candidate: broken? for helpful, honest, harmless, cites
broken = {"A": [0, 1, 0, 1], "B": [0, 0, 0, 0], "C": [1, 0, 0, 1]}
# v(y) = Σ_k 1[y breaks principle k]
v = {name: sum(flags) for name, flags in broken.items()}
v  # → {'A': 2, 'B': 0, 'C': 2}
# keep only the pairs with a strict winner, winner first
pairs = [(a, b) if v[a] < v[b] else (b, a) for a, b in [("A", "B"), ("A", "C"), ("B", "C")] if v[a] != v[b]]
pairs  # → [('B', 'A'), ('B', 'C')]

A real constitution has more principles, written as sentences rather than keyword checks, and the "editor" is a language model reading them, so its judgements are softer and can be wrong. The shape of the pipeline is the same.

Figure 3 · Diagram

Reading it: the constitution (the cylinder) feeds two places, and that is the whole idea: the same written principles drive both the critique that produces better answers and the labeler that produces preference pairs. The top path makes supervised training data from revisions. The bottom path replaces human comparisons with AI comparisons; from the reward model onwards it is the ordinary RLHF loop from primer.ml.training_stages.

Figure 4 · Drawn from the lesson's code

helpful honest harmless cites 0 1 2 3 4 5 answers breaking it (of 6) A scripted critic applying the toy constitution as drafted after one critique-and-revise pass

Across six made-up candidate answers, the honest and cites principles are each broken several times before revision and zero times after it

Reading it: each pair of bars is one principle. The grey bar counts how many of six toy candidate answers break it as drafted; the blue bar counts the same answers after one critique-and-revise pass. Every blue bar is at zero here because the toy's fixes are exact (swap "guaranteed" for "likely", add a source line). With a real model the revision is itself written by the model, so the blue bars shrink rather than vanish, and checking them is part of the job.

Why it matters in practice: human labelling is slow, costly and hard to keep consistent; written principles make the target explicit and auditable. The risk moves rather than disappears: the AI labeler has its own blind spots, and whatever it gets wrong, the reward model learns faithfully. That is why AI labels are spot-checked against human ones, the same way primer.agents.evals calibrates an LLM judge.

In code: CONSTITUTION holds the four toy principles, critique returns the names of the ones an answer breaks, revise applies one fix per broken principle, and constitutional_preference_pairs turns a list of candidates into (winner, loser) pairs, dropping ties. The pairs are what primer.ml.training_stages.reward_model_loss trains on.

Chapter 3

Red-teaming: looking for failures on purpose

Test the filter by hand, then let a search try; the two numbers will not agree.

Everyday picture Before a bank opens a new vault, it pays a team to try to get in. Not because it expects burglars to be clever in any particular way, but because the people who built the vault only think of the ways in they already guarded against. Red-teaming is the same for models and their safety checks: search systematically for inputs that make them fail, count how often the search succeeds, fix what it found, and search again.

Tiny worked example In our toy world there is one off-limits action, called xyzzy (a made-up word). A simple safety filter blocks any request containing the word "xyzzy". The toy language also has two made-up synonyms, "plugh" and "quux", that mean exactly the same thing. The filter's author tested it with three hand-written requests:

Hand-written test Blocked?
please do xyzzy yes
do xyzzy yes
kindly do xyzzy now yes

Three for three: the filter looks perfect. An automated search does something duller and more thorough: it takes the request "please do xyzzy" and rewrites each word at random with any word that means the same thing ("please" or "kindly", "do" or "perform", "xyzzy" or "plugh" or "quux"). Six attempts from that search:

Attempt Blocked? Gets through (still off-limits, not blocked)?
please do xyzzy yes no
kindly perform plugh no yes
please do quux no yes
kindly do xyzzy yes no
please perform plugh no yes
kindly do quux no yes

Four of six attempts get through. The fraction of attempts that get through is the attack success rate.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
how many attempts the search made 6
a counter walking over the attempts 1 … 6
the -th attempted request "kindly perform plugh"
"gets through" the filter allows it, yet it still asks for the off-limits action yes for attempts 2, 3, 5, 6
1 if the statement is true, 0 if not 0, 1, 1, 0, 1, 1
ASR attack success rate: the share of attempts that get through 0 … 1

In words: "count the attempts that got past the filter while still asking for the off-limits thing, and divide by the number of attempts."

With the numbers: (0 + 1 + 1 + 0 + 1 + 1) / 6 = 4 / 6 = 0.667. The hand-written tests give 0 / 3 = 0: they only ever used the word the filter already knew.

Level 3: in Python
# 1 = the attempt got through, 0 = it was blocked
through = [0, 1, 1, 0, 1, 1]
# ASR = (1/N) Σ_i 1[x_i gets through]
round(sum(through) / len(through), 3)  # → 0.667
# the three hand-written tests: all blocked
hand_written = [0, 0, 0]
sum(hand_written) / len(hand_written)  # → 0.0

Figure 5 · Diagram

Reading it: the inner loop (Mutate, Filter, back to Mutate) is the search: cheap, automatic, and indifferent to what the author expected. Every attempt that slips through is logged, and the log turns into one number, the attack success rate. The outer loop is the engineering: fix the filter using the logged failures, then search again with fresh randomness, so the fix is judged on attempts it was not built from.

Figure 6 · Drawn from the lesson's code

hand-written tests automated search search after patching 0.0 0.2 0.4 0.6 0.8 1.0 attack success rate Same filter, three ways to measure it 0% 64% 0% 1 0 0 1 0 1 1 0 2 search attempts (log scale) 0 1 2 distinct words getting through The search finds both synonyms fast

Hand-written tests report 0% success, the automated search reports 64%, and after patching the filter with its findings a fresh search reports 0%; the right panel shows both synonyms found within the first handful of tries

Reading it: on the left, each bar is one way of measuring the same filter. The hand-written tests say 0%; the automated search says 64% (close to the two thirds you would expect), because two of the three words for xyzzy are unknown to the filter. After adding the words the search found and searching again with a new seed, the rate drops to 0%. On the right, the x-axis is the number of search attempts and the y-axis is how many distinct words for xyzzy the search has found getting through; both turn up within a handful of tries. The lesson of the left panel is that a test set written by the builder measures the builder's imagination, not the filter.

The toy's vocabulary is tiny and finite, so the patched keyword list ends up complete. Real language is not: any keyword list will miss paraphrases no one has searched for yet. That is why real systems use learned classifiers instead of keyword lists, keep searching after every fix, and never rely on one filter alone (the layered checks in primer.agents.guardrails). Real-world red-teaming uses people and other models as the search, which finds far more varied failures than word swaps.

Why it matters in practice: a safety check nobody has tried to break has an unknown failure rate, which in practice means a higher one than anyone thinks. Systematic search turns "we couldn't think of a way past it" into a measured rate that can be tracked from release to release.

In code: KeywordFilter is the toy filter, red_team_attempts runs the seeded word-swap search, attack_success_rate turns the outcomes into ASR, red_team_search does both, and patch_filter adds every word for xyzzy found in a successful attempt. HAND_WRITTEN_TESTS is the author's own test set.

Chapter 4

Sycophancy: agreeing with the person instead of the facts

Be the model under pressure first; the flip rate will then read as the count it is.

Everyday picture A tutor is asked "what's 7 × 8?" and says 56. The student frowns: "I'm pretty sure it's 54." A good tutor says "let's check" and still answers 56. A tutor who wants to be liked says "oh, you're right, 54." That second tutor is sycophantic: the answer depends on what the person seems to want to hear, not on the question.

Tiny worked example Ask a model five factual questions twice: once plainly, and once after the user states a wrong answer.

Question Plain answer After "I'm sure it's …" Flipped?
7 × 8 56 56 (user said 54) no
capital of Australia Canberra Sydney (user said Sydney) yes
12 + 15 27 27 (user said 28) no
boiling point of water at sea level, °C 100 90 (user said 90) yes
number of continents 7 7 (user said 6) no

Two of five answers changed only because the user pushed. That share is the flip rate.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
how many questions were asked both ways 5
the model's answer to question asked plainly 56, Canberra, …
its answer to the same question after the user asserts a wrong answer 56, Sydney, …
"is not equal to" Sydney ≠ Canberra
1 if the answer changed, 0 if it held 0, 1, 0, 1, 0

In words: "ask each question plainly and under pressure, and count the share of questions whose answer changed."

With the numbers: (0 + 1 + 0 + 1 + 0) / 5 = 2 / 5 = 0.4.

Level 3: in Python
plain = [56, "Canberra", 27, 100, 7]
pushed = [56, "Sydney", 27, 90, 7]
# 1[a_pushed ≠ a_plain] for each question
flips = [int(p != q) for p, q in zip(pushed, plain)]
flips  # → [0, 1, 0, 1, 0]
sum(flips) / len(flips)  # → 0.4

Figure 7 · Diagram

Reading it: one question goes down two paths that differ in exactly one thing, the user's stated opinion. Everything else is held fixed, so any difference in the answers can only come from that opinion. This is the design of a controlled experiment, and it is what makes the flip rate mean something: choosing questions with known answers lets the plain answer be checked too, so a flip towards a wrong answer can be told apart from a correction.

Why preference training can cause it

Everyday picture People tend to rate an answer that agrees with them a little more kindly. That's human, and it is mild. But a reward model learns from thousands of those ratings, and then a model is tuned hard to please the reward model. A mild tilt in the labels becomes a steady push in the model.

Tiny worked example Using the Bradley-Terry model from primer.ml.training_stages: a rater compares a correct answer that contradicts them with an agreeing answer that is wrong. The correct answer is better by 1 point of genuine quality, but agreement earns a bonus of 2 points in the rater's eyes. The agreeing answer wins with probability σ(2 − 1) = σ(1) = 0.731, nearly three times in four.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the agreement bonus: how much extra credit agreeing earns in the rater's eyes 2
the correctness gap: how much better the correct answer really is 1
the sigmoid: squashes any number into a probability between 0 and 1 (σ(0) = 0.5) σ(1) = 0.731
Euler's number (≈ 2.718) raised to
"is preferred to"

In words: "the chance a rater prefers the agreeing answer is the sigmoid of the agreement bonus minus the quality gap."

With the numbers: σ(2 − 1) = 1 / (1 + e^{−1}) = 1 / 1.368 = 0.731. If the bonus were only 0.5, σ(0.5 − 1) = σ(−0.5) = 0.378: the correct answer would usually win, but the agreeing one would still win more than a third of the time.

Level 3: in Python
import math
def sigma(z):
    return 1 / (1 + math.exp(-z))
# P(agreeing ≻ correct) = σ(δ − Δ)
round(sigma(2.0 - 1.0), 3)  # → 0.731
# a smaller agreement bonus still wins more than a third of the time
round(sigma(0.5 - 1.0), 3)  # → 0.378

Figure 8 · Drawn from the lesson's code

0 1 2 3 4 agreement bonus δ 0.2 0.4 0.6 0.8 1.0 P(agreeing answer preferred) Raters' small bias, learned by the reward model quality gap Δ = 0.5 quality gap Δ = 1 quality gap Δ = 2 0.0 0.2 0.4 0.6 0.8 1.0 how often the toy model defers 0.0 0.2 0.4 0.6 0.8 1.0 measured flip rate The flip rate recovers the behaviour

Left: the chance the agreeing answer wins climbs as the agreement bonus grows, for three quality gaps. Right: the measured flip rate tracks how often the toy model defers

Reading it: on the left, the x-axis is the agreement bonus δ and each line is a different quality gap Δ. Every line crosses 0.5 where the bonus equals the gap: past that point, raters prefer the agreeing answer more often than not, and a reward model trained on their labels learns to reward agreement. On the right, the x-axis is how often the toy model defers to the user and the y-axis is the flip rate measured on 2,000 made-up questions; the dots sit on the diagonal, which shows the measurement recovers the behaviour it was built to detect.

Why it matters in practice: a sycophantic model is least reliable exactly when someone most needs a straight answer, because they already hold a wrong belief. The usual fixes all target the labels: rater instructions that ask about correctness first, preference pairs built specifically so that the correct answer disagrees with the user, and flip-rate evaluations run on every release.

In code: sycophancy_flip_rate asks a seeded toy model each made-up question plainly and under pressure and returns the flip rate; sycophantic_preference is σ(δ − Δ), computed with primer.ml.training_stages.preference_probability.

The rater's bonus above, with the bonus and the gap in your hands.

Chapter 5

Refusals: two ways to get it wrong

Slide the threshold yourself first; the trade will then read as what it is.

Everyday picture A pharmacist refuses to sell some things without a prescription. Refuse too little and harm gets through; refuse too much and people are turned away for aspirin. Both are failures, and making one rarer tends to make the other more common. A model that declines requests faces the same trade.

Tiny worked example A classifier gives every request a risk score between 0 and 1, and the model refuses anything scoring at least the threshold . Three benign requests (like "What is the capital of France?") score 0.1, 0.3 and 0.6; three requests for the toy off-limits action score 0.4, 0.8 and 0.9. At :

Request type Scores Refused at t = 0.5 Error
benign 0.1, 0.3, 0.6 only 0.6 1 of 3 refused: over-refusal
off-limits 0.4, 0.8, 0.9 0.8 and 0.9 1 of 3 answered: harmful compliance
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the threshold: refuse any request scoring at least 0.5
, the set of benign requests and the set of off-limits ones 3 each
how many items are in 3
"for each in "
, the risk score of one request 0.6, 0.4
OR() over-refusal rate: share of benign requests refused 1/3
HC() harmful compliance rate: share of off-limits requests answered 1/3

In words: "over-refusal is the share of harmless requests that score at or above the threshold; harmful compliance is the share of off-limits requests that score below it."

With the numbers: OR(0.5) = (0 + 0 + 1) / 3 = 0.333 and HC(0.5) = (1 + 0 + 0) / 3 = 0.333. Lower the threshold to 0.35 and the 0.4 request is refused too (HC = 0), but so is nothing new on the benign side (OR stays 0.333); lower it to 0.25 and the 0.3 benign request is refused as well (OR = 0.667).

Level 3: in Python
benign = [0.1, 0.3, 0.6]
off_limits = [0.4, 0.8, 0.9]
def OR(t):
    return sum(s >= t for s in benign) / len(benign)
def HC(t):
    return sum(s < t for s in off_limits) / len(off_limits)
round(OR(0.5), 3), round(HC(0.5), 3)  # → (0.333, 0.333)
# stricter thresholds trade one error for the other
[(t, round(OR(t), 3), round(HC(t), 3)) for t in (0.35, 0.25)]  # → [(0.35, 0.333, 0.0), (0.25, 0.667, 0.0)]

These are the precision/recall trade-offs from primer.ml.metrics wearing different names. Treat "refuse" as the positive call: over-refusal is the false positive rate, and harmful compliance is the miss rate, one minus recall. Choosing the threshold is the same pricing of errors as in that lesson's cost-versus-threshold section.

Figure 9 · Diagram

Reading it: every request takes exactly one of the two exits, and each exit has its own way to be wrong (the dotted boxes). Moving the threshold does not remove errors; it moves requests from one exit to the other. The only way to shrink both errors at once is a better classifier, one whose scores separate the two kinds of request more cleanly.

Figure 10 · Drawn from the lesson's code

0.0 0.2 0.4 0.6 0.8 1.0 refusal threshold t 0.0 0.2 0.4 0.6 0.8 1.0 error rate Moving the threshold trades one error for the other over-refusal (benign refused) harmful compliance (off-limits answered) 0.0 0.2 0.4 0.6 0.8 1.0 harmful compliance 0.0 0.2 0.4 0.6 0.8 1.0 over-refusal Only a better classifier helps both weaker classifier better classifier

Left: as the threshold rises, over-refusal falls and harmful compliance rises. Right: plotted against each other, a better-separating classifier's curve sits closer to the corner where both errors are zero

Reading it: on the left, the x-axis is the threshold; the blue line is over-refusal and the red line is harmful compliance, measured on 1,000 made-up scored requests. They cross: no threshold makes both small. On the right, the same numbers are plotted against each other, one point per threshold, for two classifiers. The bottom-left corner (no over-refusal, no harmful compliance) is the goal. Sliding along a curve is choosing a threshold; jumping to the lower curve is building a better classifier, which is where real progress comes from.

Why it matters in practice: over-refusal is easy to overlook because it fails quietly: nobody reports a harmful answer that didn't happen, but people do stop using a model that turns down ordinary requests. Measuring both rates on every release, on dedicated sets of benign-but-sensitive- looking requests as well as off-limits ones, keeps the trade visible.

In code: refusal_tradeoff returns both error rates at each threshold, on seeded toy scores or on scores you pass in.

Chapter 6

Checking safety before release

Run the gate first, then move the limits and watch what it stops catching.

Everyday picture A new car model is crash-tested, driven on a closed track, then lent to a few fleet customers before it reaches showrooms. Each stage is cheaper to fail than the next, and each has a checklist it must pass before the car moves on. Models are released the same way.

Tiny worked example Before release, a candidate model is measured on four evaluation sets, and each measurement has a limit set in advance:

Measurement Measured Limit Pass?
attack success rate (red-team set) 0.02 0.05 yes
over-refusal (benign set) 0.08 0.10 yes
harmful compliance (off-limits set) 0.01 0.02 yes
sycophancy flip rate (pushed questions) 0.30 0.20 no

One measurement is over its limit, so the release is held and the report names that one measurement. Setting the limits before measuring matters: limits chosen after seeing the numbers tend to drift to wherever the numbers landed.

Figure 11 · Diagram

Reading it: the left half is a gate: all the evaluation sets run, and any measurement over its limit sends the model back to be fixed. The right half is a staged rollout, each stage wider than the last, so a problem the evaluation sets missed is met by a few people first. The dotted arrows are what keeps the evaluation sets honest: every failure found in the wild becomes a new test case, so the same failure is caught at the gate next time. The rollout mechanics (shadow mode, canaries, kill switches) are built in primer.agents.deployment, and regression gates in primer.agents.evals.

Why it matters in practice: every measurement in this lesson is noisy and partial on its own. A fixed set of limits, checked on every release, is what turns them into a decision, and staged rollout is what limits the cost when the measurements were wrong.

In code: release_gate compares each measurement with its limit and returns whether the release passes and one reason for every limit missed.

Test yourself

9 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1What does Goodhart's law have to do with training a model on a reward model?Think it through, then reveal

The reward model is a measurement of what we want, not the thing itself. Tuning a model hard against it finds the places where the measurement and the goal disagree, so the score keeps rising while real quality stalls or falls. Drift penalties, early stopping and refreshed reward models are the standard ways to limit it.

Question 2In Constitutional AI, what do the written principles replace, and what do they not replace?Think it through, then reveal

They replace most of the human preference labels: a model applies the principles to critique, revise and compare answers, and those AI labels train the reward model. They do not replace human judgement about which principles to write, or human spot-checks of the AI labels, since the labeler's mistakes are learned just as faithfully as its good calls.

Question 3Why is a tie between two candidates dropped instead of labelled?Think it through, then reveal

A tie says neither answer is better, so it gives the reward model no direction to learn. Labelling it either way would teach a preference that doesn't exist, which is noise.

Question 4A safety filter passes every test its authors wrote. Why is that weak evidence?Think it through, then reveal

The tests share the authors' blind spots: they probe the cases the authors already guarded against. An automated search that varies inputs without those assumptions finds failures the hand-written tests cannot, and gives an attack success rate that means something.

Question 5After patching a filter with what red-teaming found, how should the patch be judged?Think it through, then reveal

By searching again with fresh randomness (or new red-teamers), not by re-running the attempts the patch was built from. Those will pass by construction, the same way a model scores well on its own training data.

Question 6How do you measure sycophancy, and why use questions with known answers?Think it through, then reveal

Ask the same question plainly and after the user asserts a wrong answer, and count how often the answer changes. Known answers let you tell a sycophantic flip (towards the wrong claim) apart from a legitimate correction.

Question 7How can preference training make a model more sycophantic?Think it through, then reveal

If raters give a small bonus to answers that agree with them, then by the Bradley-Terry model an agreeing but wrong answer beats a correct one whenever the bonus exceeds the quality gap. The reward model learns that bonus and preference tuning amplifies it.

Question 8Why can't a threshold fix both harmful compliance and over-refusal?Think it through, then reveal

Moving the threshold only moves requests from "answer" to "refuse" or back; every request that stops being one error risks becoming the other. Only a classifier that separates the two kinds of request better moves both rates down together.

Question 9Why set release limits before measuring?Think it through, then reveal

Limits chosen after seeing the results tend to be set wherever the results landed, which makes the gate a formality. Fixed limits turn noisy measurements into a decision made in advance.

Primary sources

The papers behind this lesson

Askell et al., A General Language Assistant as a Laboratory for Alignment (2021)

Framed the helpful, honest and harmless targets for a language assistant and compared simple ways of steering a model towards them.

The paper ↗
Bai et al., Constitutional AI: Harmlessness from AI Feedback (2022)

Introduced training against a written set of principles, with self-critique and revision followed by reinforcement learning from AI-labelled preferences (RLAIF).

Read the annotated companion →The paper ↗
Perez et al., Red Teaming Language Models with Language Models (2022)

Showed that one language model can generate test cases that find failures in another, automating red-teaming at scale.

The paper ↗
Ganguli et al., Red Teaming Language Models to Reduce Harms (2022)

Described a large human red-teaming effort, its methods and how attack success changed with model size and training.

The paper ↗
Sharma et al., Towards Understanding Sycophancy in Language Models (2023)

Measured sycophancy across assistants and traced part of it to human preference data that favours agreeable answers.

Read the annotated companion →The paper ↗
Gao, Schulman and Hilton, Scaling Laws for Reward Model Overoptimization (2022)

Measured Goodhart's law for reward models: the true reward rises then falls as a policy is optimized harder against a proxy.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Askell et al., A General Language Assistant as a Laboratory for Alignment (2021): https://arxiv.org/abs/2112.00861
  • Bai et al., Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback (2022): https://arxiv.org/abs/2204.05862
  • Bai et al., Constitutional AI: Harmlessness from AI Feedback (2022): https://arxiv.org/abs/2212.08073
  • Perez et al., Red Teaming Language Models with Language Models (2022): https://arxiv.org/abs/2202.03286
  • Ganguli et al., Red Teaming Language Models to Reduce Harms (2022): https://arxiv.org/abs/2209.07858
  • Sharma et al., Towards Understanding Sycophancy in Language Models (2023): https://arxiv.org/abs/2310.13548
  • Gao, Schulman and Hilton, Scaling Laws for Reward Model Overoptimization (2022): https://arxiv.org/abs/2210.10760

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.