rumblr Work in progressWIP

● The AI Primer · Lesson 20 · Part 1: how the model works inside

Metrics

how you know whether a model is any good

You'll be able to explain Precision/recall/F1, ROC-AUC, recall@k, MRR, nDCG, BLEU/ROUGE

Members · open during launch 32 min9 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Precision: of what I flagged, how much was right. Recall: of what mattered, how much I found. F1 balances them.
  2. Accuracy lies on imbalanced data. Flagging nothing on 1% fraud scores 99%.
  3. ROC-AUC is the probability a random positive outranks a random negative.
  4. Choose the threshold by what each kind of error costs.
  5. For RAG, measure recall@k first. What isn't retrieved can't be used.
  6. BLEU/ROUGE punish correct paraphrases; use an LLM judge and calibrate it against humans with Cohen's kappa.

Level 1

The practitioner's guide

In one sentence

A metric is the number you use to decide whether a model is good enough, as opposed to the loss the model optimises (primer.ml.losses), and every metric answers exactly one question while staying silent on all the others.

When you need it

Every time a number decides something: whether a model ships, which of two retrievers your RAG system keeps, whether an automated judge may replace human review. The tell that you have the wrong metric is a number that looks great offline and a system that disappoints. This lesson's fraud detector has 94% accuracy and misses 4 frauds in 10; on traffic that is 1% fraud, a model that flags nothing scores 99% accuracy with recall 0. Its generation example is starker: against the reference "the meeting was moved to Friday because the manager is sick", the answer that says Monday scores BLEU 0.73 and ROUGE-L 0.91, and the correct paraphrase scores 0.16 and 0.26. And a judge that stamps "pass" on everything agrees with humans 90% of the time on a sample that is 90% passes, with a Cohen's kappa of exactly 0. You do not need this lesson to read a training curve; that is the loss. You need it the moment a number leaves the training loop and enters a decision.

Your options

By the question they answer:

Metric The question it answers What it hides What it needs Where it lives
Accuracy Of all decisions, what share were right? Class imbalance: flag nothing on 1% fraud and score 99% Labels and a threshold Balanced classification only
Precision, recall, F1 Of what I flagged, how much was right; of what mattered, how much I found; F1 balances the two The threshold that produced them, and that F1 prices a miss and a false alarm equally Labels and a threshold Every fraud, spam and moderation filter
ROC-AUC Across every threshold, how often does a random positive outrank a random negative? On rare events, false alarms can swamp the true positives while the false-positive rate stays tiny Scores and labels Comparing models before a threshold is chosen
Precision-recall curve At each level of recall, what share of the flags are real? Nothing about ranking below the recall you care about Scores and labels Rare-event problems
Cost-weighted threshold Which cut-off makes misses times their price plus false alarms times theirs smallest? It needs the prices, which only the business knows Two prices The decision itself
recall@k, precision@k, MRR, nDCG@k Did we find the relevant documents, how much noise came with them, how high is the first hit, are the best ones on top? Anything the golden set does not cover; nDCG needs graded relevance A golden set: queries with the ids of the documents that answer them Search and RAG retrieval
BLEU, ROUGE-L How much wording does the answer share with a reference? Truth: a correct paraphrase scores near zero One or more reference answers Machine translation, summarisation, legacy pipelines
Embedding similarity (BERTScore) How close in meaning is the answer to a reference? It still needs a reference, and closeness is not correctness A reference and an embedding model Generation with references
LLM judge, calibrated with Cohen's kappa Does a rubric-following model agree with human graders beyond chance? The judge's own biases (position, verbosity, self-enhancement) and chance agreement A human-labelled sample and a rubric Open-ended generation and agent evaluation

How to choose

Name the question first, then the metric that answers it.

  • A classifier with a threshold in production: report precision and recall at that threshold, and pick the threshold by pricing the two errors. In this lesson, pricing a miss at 10 and a false alarm at 1 drops the threshold and lifts recall from 0.30 to 0.88; the reverse prices raise it until precision is 1.00. Never lead with accuracy on imbalanced data.
  • Comparing models before any threshold exists: ROC-AUC. If positives are rare, look at the precision-recall curve as well.
  • Retrieval, including the retrieval half of RAG: recall@k with k set to the number of chunks you actually pass to the model, because the model cannot use what was not retrieved. Add MRR or nDCG when position matters.
  • Generation: a code-based check wherever the answer can be verified (a test passes, a number matches, the JSON parses). Overlap metrics only where the task is nearly verbatim, such as translation. Everything else, an LLM judge with a written rubric, calibrated against people with kappa before it grades anything at scale.
  • Whatever you pick, a metric is meaningless without its conditions: the threshold, the k, the golden set, the sample the judge was checked on. Report them next to the number, every time.

What it costs

Metrics cost labels, not compute. Classification needs labelled examples and, for the threshold, two prices you must extract from whoever owns the consequences. Retrieval needs a golden set of real queries paired with the documents that answer them, which is hours of a knowledgeable person's time and the single best investment in a RAG system, because you can then rerun recall@k after every change to chunking, the embedding model or the reranker. Generation costs the most: reference answers for overlap metrics, or human labels for the sample a judge is calibrated on, plus a model call per graded output for the judge itself, and the calibration is repeated every time the judge model or the rubric changes. The cost of the wrong metric is the one that matters: it is the gap between the dashboard's 94% and the four frauds in ten that walked through.

What breaks

  • Accuracy on imbalanced data. 99% for a model that does nothing. Use precision and recall.
  • A number with no threshold. "What's the accuracy?" is incomplete without "at what threshold?"; every classification metric changes when the cut-off moves.
  • AUC on rare events. This lesson's classifier has an AUC of 0.89 and, at 80% recall, only about a third of its flags are real against a 10% base rate. Read the precision-recall curve.
  • F1 when the errors cost differently. The harmonic mean pulls toward the smaller of the two: precision 1.0 with recall 0.1 gives F1 0.18. If a miss costs ten times a false alarm, F1 is the wrong target; price the errors instead.
  • recall@k at the wrong k. Recall@20 is no comfort when you pass five chunks to the model. Measure at the k you serve.
  • Overlap metrics on paraphrase. The Monday answer, factually wrong, scores 0.91 on ROUGE-L. Overlap measures wording, not truth.
  • Raw agreement for a judge. 90% agreement and kappa 0 describe the same lazy judge. Kappa above about 0.6 is usually read as substantial agreement; below it, fix the rubric, the examples or the judge model.
  • A judge that drifts. A new judge model or an edited rubric is a new judge. Re-run the calibration.

In the wild

scikit-learn ships the classification family as precision_score, recall_score, f1_score, roc_auc_score, average_precision_score and cohen_kappa_score, and its metrics guide is in Further reading. In retrieval, the BEIR benchmark (Thakur et al., 2021) compares lexical, sparse, dense, late-interaction and reranking systems across 18 datasets. BLEU (Papineni et al., 2002) and ROUGE (Lin, 2004) are still the reported numbers in translation and summarisation, with BERTScore (Zhang et al., 2019) as the embedding-based successor. For judging, Zheng et al. (2023) found that strong LLM judges reach over 80% agreement with humans on MT-Bench, the level humans reach with each other, and named the position, verbosity and self-enhancement biases every judge pipeline now guards against. primer.agents.evals turns these metrics into a release gate for an agent.

Go deeper

Level 2 builds each family from a hand-sized example: the four cells of the confusion matrix and the accuracy trap, the ROC curve as a walk whose area counts correctly ordered pairs, the cheapest threshold under two prices, the four retrieval metrics on one five-document list, BLEU and ROUGE-L on the Monday sentence, and Cohen's kappa on twenty essays. If you only needed to pick the metric and read it honestly, you are done.

Level 2

How it works, from scratch

A loss function is what the model optimizes during training; a metric is what you use to decide whether it's working. Picking the wrong metric is one of the most common ways a project looks great offline and fails in production. This lesson builds every metric from scratch in three families:

  1. Classification: precision, recall, F1, accuracy, the confusion matrix, ROC-AUC, precision-recall curves, and choosing a threshold by the cost of each kind of error.
  2. Retrieval: recall@k, precision@k, MRR and nDCG, the metrics for search and retrieval-augmented generation (RAG).
  3. Generation: BLEU and ROUGE-L (word overlap), why they fail on language-model output, and how to check an automated judge against people with Cohen's kappa.

Chapter 1

Classification: precision, recall and the confusion matrix

Everyday picture You're fishing for trout in a lake that also holds old boots. Precision asks: of everything in your net, how much is trout? Recall asks: of all the trout in the lake, how many did you catch? A tiny net held in one good spot has high precision and low recall; dragging a huge net across the whole lake has high recall and a lot of boots.

Tiny worked example: fraud detection. Of 100 transactions, 10 are fraud. The model flags 8, and 6 of those are truly fraud. So 6 hits (true positives, TP), 2 false alarms (false positives, FP), 4 misses (false negatives, FN) and 88 correct passes (true negatives, TN).

Metric Calculation Result
Precision 6 correct ÷ 8 flagged 75%
Recall 6 found ÷ 10 actual fraud 60%
F1 2 × 0.75 × 0.60 ÷ (0.75 + 0.60) 67%
Accuracy (6 + 88 correct) ÷ 100 94%

Accuracy looks great at 94% while the model misses 4 in 10 frauds.

Figure 1 · Diagram

Reading it: a classifier never outputs "fraud" directly; it outputs a score, and a threshold that you choose turns scores into decisions. Comparing each decision with the true label drops the item into one of four cells of the confusion matrix, and every classification metric is just a ratio of those four counts. Change the threshold and every metric changes, which is why "what's the accuracy?" is incomplete without "at what threshold?".
predicted positive predicted negative
actual positive TP = 6 FN = 4 (missed)
actual negative FP = 2 (false alarm) TN = 88
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
TP, FP, FN, TN counts of hits, false alarms, misses and correct passes whole numbers
, precision and recall 0 … 1
the harmonic mean of P and R: an average that is dragged towards the smaller of the two 0 … 1
total number of items, TP + FP + FN + TN whole number
fraction bar "divided by"

In words: precision is hits over everything flagged; recall is hits over everything that should have been flagged; F1 is twice their product over their sum; accuracy is everything right over everything.

On the worked example: P = 6/(6 + 2) = 0.75; R = 6/(6 + 4) = 0.60; F1 = 2 × 0.75 × 0.60 / 1.35 = 0.667; accuracy = (6 + 88)/100 = 0.94. The harmonic mean punishes imbalance: P = 1.0 with R = 0.1 gives F1 = 0.18, not the ordinary average of 0.55.

Level 3: in Python
TP, FP, FN, TN = 6, 2, 4, 88
N = TP + FP + FN + TN
# precision, recall
P, R = TP / (TP + FP), TP / (TP + FN)
P, R  # → (0.75, 0.6)
# F1, accuracy
round(2 * P * R / (P + R), 3), (TP + TN) / N  # → (0.667, 0.94)
# F1 when P = 1.0 and R = 0.1
round(2 * 1.0 * 0.1 / (1.0 + 0.1), 2)  # → 0.18

In code: confusion sorts labels and decisions into a Confusion, whose Confusion.precision, Confusion.recall, Confusion.f1 and Confusion.accuracy are the four formulas. fraud_example rebuilds the 100 transactions above, and flag_nothing_trap builds the model below that never flags anything.

Why it matters in practice: the accuracy trap. If 1% of transactions are fraud, a model that flags nothing is 99% accurate and completely useless (recall 0). On imbalanced data, never lead with accuracy.

Chapter 2

ROC-AUC: how well does the model rank?

Everyday picture A smoke alarm has a sensitivity dial. Turn it up and it catches every fire but also shrieks at toast; turn it down and it stays quiet but might miss a real fire. The ROC curve draws every dial setting at once, and AUC scores the alarm across all of them.

Tiny worked example Two frauds scored 0.9 and 0.4, two legitimate transactions scored 0.6 and 0.1. Compare every (fraud, legitimate) pair: 0.9 > 0.6 ✓, 0.9 > 0.1 ✓, 0.4 > 0.6 ✗, 0.4 > 0.1 ✓. Three of four pairs are ranked correctly, so AUC = 0.75.

Figure 4 · Diagram

Reading it: the ROC curve is a walk. Starting from "flag nothing" (bottom-left) you lower the threshold one score at a time; each positive you pick up moves you up, each negative moves you right. A perfect ranker goes straight up then straight right (area 1.0); a random one wanders along the diagonal (area 0.5). Because the walk moves up exactly when a positive outranks the remaining negatives, the area counts correctly ordered (positive, negative) pairs, which is why AUC equals the pairwise win rate. The code computes AUC both ways (trapezoids under the curve, and counting pairs) and they agree exactly.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
TPR true-positive rate: recall, the share of positives flagged 0 … 1
FPR false-positive rate: the share of negatives wrongly flagged 0 … 1
, (in the AUC sum) the set of positive items and the set of negative items sets
, how many items each set holds whole numbers
, the model's score for positive p and negative n real
1 if the statement inside is true, else 0 0 or 1
add over every (positive, negative) pair

In words: AUC is the share of (positive, negative) pairs in which the positive gets the higher score, counting ties as half.

On the worked example: 2 positives × 2 negatives = 4 pairs; 3 are ordered correctly; AUC = 3/4 = 0.75. One point on the curve: at threshold 0.5 the model flags the 0.9 fraud and the 0.6 legitimate transaction, so TPR = 1/(1 + 1) = 0.5 and FPR = 1/(1 + 1) = 0.5.

Level 3: in Python
# scores of the positives (frauds)
s_P = [0.9, 0.4]
# scores of the negatives (legitimate)
s_N = [0.6, 0.1]
t = 0.5
TP, FN = sum(s >= t for s in s_P), sum(s < t for s in s_P)
FP, TN = sum(s >= t for s in s_N), sum(s < t for s in s_N)
# TPR, FPR
TP / (TP + FN), FP / (FP + TN)  # → (0.5, 0.5)
pairs = sum((s_p > s_n) + 0.5 * (s_p == s_n) for s_p in s_P for s_n in s_N)
# AUC: the share of pairs ranked correctly
pairs / (len(s_P) * len(s_N))  # → 0.75

Figure 2 · Drawn from the lesson's code

0.0 0.2 0.4 0.6 0.8 1.0 false-positive rate (share of negatives flagged) 0.0 0.2 0.4 0.6 0.8 1.0 true-positive rate (recall) ROC curve model (AUC = 0.89) random (AUC = 0.50)

ROC curve bowing well above the coin-flip diagonal: the classifier ranks a random positive above a random negative 89% of the time (AUC 0.89)

Reading it: the solid line is the ROC curve of synthetic_scores() (positives shifted 1.5 standard deviations above negatives); the dashed diagonal is a coin flip. The shaded area is the AUC, about 0.89: pick one fraud and one legitimate transaction at random and the model scores the fraud higher 89% of the time. Where the curve bends is where a sensible threshold lives.

Figure 3 · Drawn from the lesson's code

0.0 0.2 0.4 0.6 0.8 1.0 recall 0.0 0.2 0.4 0.6 0.8 1.0 precision Precision-recall curve model base rate (10% positive)

Precision-recall curve falling as recall rises: at 80% recall only about a third of the flags are real, against a 10% base rate

Reading it: the same scores, viewed as precision (y) against recall (x). Moving right means lowering the threshold: you find more of the positives, but precision falls as false alarms pile in. The dotted line is the base rate (10% positives), i.e. what flagging at random achieves. On rare-event problems this plot tells the truth that ROC's tiny false-positive rates hide: at 80% recall only about a third of the flags are real.

In code: roc_curve takes the walk, roc_auc measures the area under it with auc_trapezoid, and roc_auc_rank counts correctly ordered pairs instead; Confusion.fpr is FPR, and pr_curve gives the precision-recall points.

Why it matters in practice. AUC compares models independently of any threshold. On heavily imbalanced data, look at the precision-recall curve too, because FPR stays tiny even when false alarms swamp the true positives.

Chapter 3

Picking the threshold: price your errors

Everyday picture A hospital screening test and a spam filter want opposite things. Missing a disease is terrible, so the screen flags generously; losing an important email is annoying, so the spam filter flags cautiously. Same maths, different prices.

Tiny worked example With the four scores above (frauds 0.9 and 0.4, legitimate 0.6 and 0.1): if a miss costs 10 and a false alarm costs 1, the cheapest rule is "flag at 0.4 or above", catching both frauds and wrongly flagging one legitimate transaction (total cost 1). If a false alarm costs 10 and a miss 1, the cheapest rule is "flag only 0.9", missing one fraud but raising no false alarms (total cost 1).

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
the threshold: flag items scoring at least t real
, misses and false alarms at that threshold whole numbers
, the price of one miss and of one false alarm ≥ 0
"the value of t that makes this smallest"
the best threshold real

In words: the total cost at a threshold is misses times their price plus false alarms times theirs; pick the threshold where that total is lowest.

On the worked example: c_FN = 10, c_FP = 1: at t = 0.4, FN = 0 and FP = 1, cost 1, the minimum. The other thresholds cost 10 (t = 0.9, one miss), 11 (t = 0.6, one miss and one false alarm) and 2 (t = 0.1, two false alarms).

Level 3: in Python
frauds, legit = [0.9, 0.4], [0.6, 0.1]
c_FN, c_FP = 10, 1
def cost(t):
    # frauds below t are missed
    FN = sum(s < t for s in frauds)
    # legitimate ones at or above t are false alarms
    FP = sum(s >= t for s in legit)
    return c_FN * FN + c_FP * FP
[cost(t) for t in (0.9, 0.6, 0.4, 0.1)]  # → [10, 11, 1, 2]
# t* = argmin_t cost(t)
min((0.9, 0.6, 0.4, 0.1), key=cost)  # → 0.4

Figure 5 · Drawn from the lesson's code

−4 −3 −2 −1 0 1 2 3 threshold (flag if score >= threshold) 0 100 200 300 400 500 total cost Where to cut depends on what errors cost miss costs 1, false alarm costs 1 miss costs 10, false alarm costs 1 miss costs 1, false alarm costs 10

Total cost against threshold for three error prices: the cheapest threshold moves left when misses cost more and right when false alarms cost more

Reading it: each line is the total cost (misses × their price + false alarms × their price) at every threshold, and each dot marks that line's minimum. When a miss costs 10× a false alarm, the best threshold slides left (flag more, higher recall); when a false alarm costs 10×, it slides right (flag less, higher precision). The model is the same in all three lines; only the business decides where to cut.

In code: best_threshold_by_cost tries every distinct score as t (plus "flag nothing") and keeps the cheapest.

Chapter 4

Retrieval metrics: judging a search result list

Everyday picture You ask a librarian for books on a topic and get a stack of five. Did the stack include the books that matter (recall)? How much of the stack is useful (precision)? Is a good book on top, or do you dig for it (reciprocal rank)? Are the best books nearest the top (nDCG)?

Tiny worked example One query. The system returns d7, d3, d9, d1, d4 in that order. The relevant documents are d3 (very relevant, grade 3), d4 (grade 2) and d8 (grade 1, never returned).

  • recall@5 = 2 found of 3 relevant = 0.667
  • precision@5 = 2 relevant of 5 returned = 0.4
  • reciprocal rank = first relevant result at rank 2, so 1/2 = 0.5
  • nDCG@5 = 0.560 (worked below)

Figure 7 · Diagram

Reading it: retrieval is evaluated separately from generation. A golden set pairs real queries with the ids of the documents that answer them; the retriever produces a ranked list for each query; four metrics ask four different questions of the same list. Averaging them over the golden set gives numbers you can track every time you change chunking, the embedding model or the reranker.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
how many top results you look at e.g. 5
"in both": documents that are relevant and in the top k set
how many items a set holds
, the set of test queries, and one query
position (1 = top) of the first relevant result for query q 1, 2, …
the relevance grade of the result at position i (0 if irrelevant) e.g. 0 … 3
logarithm base 2: how many times you halve i + 1 to reach 1. It grows slowly, so it is a gentle position discount 1 at rank 1, 1.58 at rank 2
IDCG "ideal DCG": the DCG of the same grades sorted best-first, so nDCG tops out at 1

In words: recall@k is the share of relevant documents found in the top k; MRR averages one-over-the-rank of the first hit; DCG adds each result's grade, discounted by the log of its position; nDCG divides by the best possible DCG.

On the worked example: recall@5 = 2/3 = 0.667; with one query whose first hit is at rank 2, MRR = 1/2 = 0.5. DCG = 3/log₂(3) + 2/log₂(6) = 1.8928 + 0.7737 = 2.6665 (d3 at rank 2, d4 at rank 5). Ideal order d3, d4, d8: IDCG = 3/1 + 2/1.585 + 1/2 = 4.7619. nDCG = 2.6665/4.7619 = 0.560.

Level 3: in Python
import math
ranked = ["d7", "d3", "d9", "d1", "d4"]
# relevance grades; missing means 0
rel = {"d3": 3, "d4": 2, "d8": 1}
k = 5
# recall@k
round(len(set(rel) & set(ranked[:k])) / len(rel), 3)  # → 0.667
rank_q = next(i for i, doc in enumerate(ranked, start=1) if doc in rel)
# MRR over a single query
1 / rank_q  # → 0.5
DCG = sum(rel.get(doc, 0) / math.log2(i + 1) for i, doc in enumerate(ranked[:k], start=1))
# the same grades, best first
best = sorted(rel.values(), reverse=True)[:k]
IDCG = sum(g / math.log2(i + 1) for i, g in enumerate(best, start=1))
print(f"{DCG:.4f} {IDCG:.4f} {DCG / IDCG:.3f}")  # → 2.6665 4.7619 0.560

Figure 6 · Drawn from the lesson's code

1 2 3 4 5 6 7 8 9 10 rank position 0.0 0.2 0.4 0.6 0.8 1.0 weight 1 / log2(rank + 1) How much each position counts in DCG

DCG weight by rank: rank 1 counts fully, rank 2 counts 63% and rank 10 still counts 29%

Reading it: the bars are the weight 1/log2(rank + 1) that DCG gives a result at each position. Rank 1 counts fully, rank 2 counts 63%, rank 10 only 29%. That gentle logarithmic decay says "position matters, but a good result at rank 5 is still worth a lot", which matches how people and language models actually use a result list.

In code: recall_at_k and precision_at_k score one ranked list, reciprocal_rank finds its first hit and mean_reciprocal_rank averages that over queries; dcg_at_k adds the discounted grades and ndcg_at_k divides by the ideal ordering's DCG.

Why it matters in practice. For RAG, recall@k (with k = the number of chunks you pass to the model) usually matters most: the model can't use a passage that wasn't retrieved, and no prompt fixes that.

Chapter 5

Generation metrics: why word overlap misleads

Everyday picture Grading essays by counting how many words each one shares with the answer key. A student who copies the key's phrasing but gets the key fact wrong scores well; a student who explains it correctly in their own words scores badly.

Tiny worked example Reference: "The meeting was moved to Friday because the manager is sick."

candidate BLEU ROUGE-L
"Since the boss is ill, the team rescheduled the meeting for Friday." (correct) 0.16 0.26
"The meeting was moved to Monday because the manager is sick." (wrong) 0.73 0.91

BLEU (from machine translation) multiplies together how many of the candidate's 1-, 2-, 3- and 4-word sequences appear in the reference, with a penalty for being too short. ROUGE-L (from summarisation) measures the longest sequence of words the two share in the same order, gaps allowed.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
share of the candidate's n-word sequences also found in the reference, each counted at most as often as the reference has it 0 … 1
, natural log and its inverse; exp of the average log is the geometric mean of the four
BP brevity penalty: 1 if the candidate is longer than the reference, else e^(1 − ref length / candidate length) 0 … 1
LCS longest common subsequence: the longest run of words appearing in both texts in the same order
, LCS length over candidate length, and over reference length 0 … 1

In words: BLEU is the geometric mean of the n-gram match rates, scaled down for short answers; ROUGE-L is the F1 of the longest shared in-order word sequence.

On a hand example: candidate "a b c d", reference "a c d e": the LCS is "a c d", 3 words, so P = 3/4, R = 3/4, ROUGE-L = 0.75. And "the the the the" against "the cat is here" scores unigram precision 1/4: "the" is credited only as often as the reference contains it.

With the numbers: the Monday answer and the reference are both 11 words (full stop and capitals dropped), so BP = 1. It matches 10 of its 11 words, 8 of its 10 word pairs, 6 of its 9 triples and 4 of its 8 four-word runs. bleu adds 1 to the top and bottom for n ≥ 2 (smoothing, so one missing four-word run can't zero the score), giving p = 10/11, 9/11, 7/10, 5/9, and BLEU = exp(¼(ln 0.909 + ln 0.818 + ln 0.7 + ln 0.556)) = 0.73. Its longest shared in-order run is 10 words, so ROUGE-L = 2 · (10/11) · (10/11) / (20/11) = 0.91.

In Python:

import math
ref = "the meeting was moved to friday because the manager is sick".split()
cand = "the meeting was moved to monday because the manager is sick".split()
# every run of n words, in order
def grams(words, n):
    return [tuple(words[i:i + n]) for i in range(len(words) - n + 1)]
def p(n):
    c, r = grams(cand, n), grams(ref, n)
    # clipped: at most as often as ref has it
    hits = sum(min(c.count(g), r.count(g)) for g in set(c))
    # add-one smoothing for n ≥ 2
    s = 1 if n > 1 else 0
    return (hits + s) / (len(c) + s)
[round(p(n), 3) for n in range(1, 5)]  # → [0.909, 0.818, 0.7, 0.556]
BP = 1.0 if len(cand) > len(ref) else math.exp(1 - len(ref) / len(cand))
# BLEU
round(BP * math.exp(sum(math.log(p(n)) for n in range(1, 5)) / 4), 2)  # → 0.73
# every word but "monday", in order
P_LCS, R_LCS = 10 / len(cand), 10 / len(ref)
# ROUGE-L
round(2 * P_LCS * R_LCS / (P_LCS + R_LCS), 2)  # → 0.91
# the hand example: LCS "a c d"
P_LCS, R_LCS = 3 / 4, 3 / 4
2 * P_LCS * R_LCS / (P_LCS + R_LCS)  # → 0.75

Figure 8 · Drawn from the lesson's code

correct paraphrase wrong but copies wording exact copy 0.0 0.2 0.4 0.6 0.8 1.0 score Overlap metrics reward wording, not truth BLEU ROUGE-L

BLEU and ROUGE-L bars: the correct paraphrase scores lowest while the wrong answer that copies the wording scores almost as high as an exact copy

Reading it: three candidate answers to the same reference ("the meeting was moved to Friday because the manager is sick"). The correct paraphrase (left) scores lowest on both metrics; the answer that says Monday (middle) scores nearly as high as an exact copy (right). Overlap metrics reward wording, not truth.

In code: bleu clips each n-gram count, takes the geometric mean and applies the brevity penalty; rouge_l finds the longest common subsequence and returns its F1.

Why it matters in practice. Teams grade open-ended output with an LLM as a judge, a model following an explicit rubric, and embedding-based scores like BERTScore do somewhat better than overlap. But a judge must earn trust first (next section).

Chapter 6

Calibrating a judge: Cohen's kappa

Everyday picture Two teachers grade the same 20 essays pass/fail and agree on 18. Impressive? Not if 18 of the essays were obvious passes: a careful teacher who reads every essay (and passes those 18) and a lazy one who stamps "pass" on everything without reading would also agree on 18. Cohen's kappa subtracts the agreement you'd expect from luck.

Tiny worked example Human labels: 18 pass, 2 fail. A lazy judge says "pass" to everything. Raw agreement is 90%, and chance agreement (both say pass at their own rates) is also 0.9 × 1.0 + 0.1 × 0 = 90%, so kappa = 0. A second example: labels (y, y, n, n) vs. (y, n, n, n) agree 3/4 of the time; chance predicts 0.5 × 0.25 + 0.5 × 0.75 = 0.5; kappa = 0.5.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
kappa: agreement beyond chance ≤ 1; 0 = chance, 1 = perfect
observed agreement: share of items both raters labelled the same 0 … 1
agreement expected by chance 0 … 1
one possible label (pass, fail…)
, how often rater A (and rater B) uses label ℓ 0 … 1

In words: kappa is how far agreement beats chance, as a share of the most it could beat chance by.

On the worked example: p_o = 0.75, p_e = 0.5, κ = (0.75 − 0.5)/(1 − 0.5) = 0.5.

Level 3: in Python
# rater A's labels
A = ["y", "y", "n", "n"]
# rater B's labels
B = ["y", "n", "n", "n"]
p_o = sum(a == b for a, b in zip(A, B)) / len(A)
# Σ_ℓ p_A(ℓ) p_B(ℓ)
p_e = sum((A.count(ell) / len(A)) * (B.count(ell) / len(B)) for ell in ("y", "n"))
p_o, p_e  # → (0.75, 0.5)
# κ
(p_o - p_e) / (1 - p_e)  # → 0.5

Figure 9 · Diagram

Reading it: the judge earns trust the same way a new human reviewer would: grade a sample that people have already graded and compare. Kappa, not raw agreement, is the gate, because on a sample that is mostly "pass" a lazy judge agrees by accident. Only once the judge clears the bar do you let it grade thousands of outputs, and you repeat the check whenever the judge model or rubric changes.

In code: cohens_kappa computes p_o and p_e from two lists of labels and returns κ.

Why it matters in practice. Kappa above about 0.6 is usually considered substantial agreement. Re-check it whenever the judge model or rubric changes.

Test yourself

6 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: A fraud model has 94% accuracy. Is it good?Think it through, then reveal

A: You can't tell from accuracy. If fraud is 10% of traffic, check recall and precision. In the worked example recall is only 60%, so 4 in 10 frauds slip through.

Question 2Q: When do you optimize for precision vs. recall?Think it through, then reveal

A: By the cost of each error. If a miss is expensive (fraud, cancer screening, a relevant legal document), favor recall. If a false alarm is expensive (blocking a good customer, paging an engineer at 3 a.m.), favor precision. Set the threshold to minimize expected cost.

Question 3Q: What does an AUC of 0.8 mean?Think it through, then reveal

A: A randomly chosen positive gets a higher score than a randomly chosen negative 80% of the time. It measures ranking quality across all thresholds; 0.5 is random.

Question 4Q: Which retrieval metric matters most for RAG, and why?Think it through, then reveal

A: Recall@k, where k is the number of chunks you pass to the model. If the right passage isn't in the top k, no prompt engineering can recover the answer. Add MRR or nDCG when position within the context matters.

Question 5Q: Why not grade LLM answers with BLEU or ROUGE?Think it through, then reveal

A: They measure surface overlap with one reference. A correct paraphrase scores low and a wrong answer that reuses the reference's words scores high. Use code-based checks where possible and a rubric-driven LLM judge otherwise, validated against human labels.

Question 6Q: Your LLM judge agrees with humans 90% of the time. Good enough?Think it through, then reveal

A: Not necessarily. If 90% of answers are "pass", a judge that always says "pass" also agrees 90%. Compute Cohen's kappa to correct for chance agreement.

Primary sources

The papers behind this lesson

Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023): Measured how well strong language models agree with human preferences when used as judges, and catalogued their biases (position, verbosity, self-preference).

Read the annotated companion →The paper ↗

Papineni et al., BLEU: a Method for Automatic Evaluation of Machine Translation (2002): Introduced clipped n-gram precision with a brevity penalty.

The paper ↗

Lin, ROUGE: A Package for Automatic Evaluation of Summaries (2004): Introduced recall-oriented overlap measures, including the LCS-based ROUGE-L.

The paper ↗

Zhang et al., BERTScore: Evaluating Text Generation with BERT (2019): Compared candidate and reference by embedding similarity rather than exact word overlap.

The paper ↗

Researcher's shelf

Further reading

  • scikit-learn, Metrics and scoring: https://scikit-learn.org/stable/modules/model_evaluation.html
  • Receiver operating characteristic (Wikipedia): https://en.wikipedia.org/wiki/Receiver_operating_characteristic
  • Discounted cumulative gain (Wikipedia): https://en.wikipedia.org/wiki/Discounted_cumulative_gain
  • Cohen's kappa (Wikipedia): https://en.wikipedia.org/wiki/Cohen%27s_kappa

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.