The lesson in one minute
What you'll be able to explain
- Precision: of what I flagged, how much was right. Recall: of what mattered, how much I found. F1 balances them.
- Accuracy lies on imbalanced data. Flagging nothing on 1% fraud scores 99%.
- ROC-AUC is the probability a random positive outranks a random negative.
- Choose the threshold by what each kind of error costs.
- For RAG, measure recall@k first. What isn't retrieved can't be used.
- BLEU/ROUGE punish correct paraphrases; use an LLM judge and calibrate it against humans with Cohen's kappa.
Level 1
The practitioner's guide
In one sentence
A metric is the number you use to decide whether a
model is good enough, as opposed to the loss the model optimises
(primer.ml.losses), and every metric answers exactly one question while
staying silent on all the others.
When you need it
Every time a number decides something: whether a model ships, which of two retrievers your RAG system keeps, whether an automated judge may replace human review. The tell that you have the wrong metric is a number that looks great offline and a system that disappoints. This lesson's fraud detector has 94% accuracy and misses 4 frauds in 10; on traffic that is 1% fraud, a model that flags nothing scores 99% accuracy with recall 0. Its generation example is starker: against the reference "the meeting was moved to Friday because the manager is sick", the answer that says Monday scores BLEU 0.73 and ROUGE-L 0.91, and the correct paraphrase scores 0.16 and 0.26. And a judge that stamps "pass" on everything agrees with humans 90% of the time on a sample that is 90% passes, with a Cohen's kappa of exactly 0. You do not need this lesson to read a training curve; that is the loss. You need it the moment a number leaves the training loop and enters a decision.
Your options
By the question they answer:
| Metric | The question it answers | What it hides | What it needs | Where it lives |
|---|---|---|---|---|
| Accuracy | Of all decisions, what share were right? | Class imbalance: flag nothing on 1% fraud and score 99% | Labels and a threshold | Balanced classification only |
| Precision, recall, F1 | Of what I flagged, how much was right; of what mattered, how much I found; F1 balances the two | The threshold that produced them, and that F1 prices a miss and a false alarm equally | Labels and a threshold | Every fraud, spam and moderation filter |
| ROC-AUC | Across every threshold, how often does a random positive outrank a random negative? | On rare events, false alarms can swamp the true positives while the false-positive rate stays tiny | Scores and labels | Comparing models before a threshold is chosen |
| Precision-recall curve | At each level of recall, what share of the flags are real? | Nothing about ranking below the recall you care about | Scores and labels | Rare-event problems |
| Cost-weighted threshold | Which cut-off makes misses times their price plus false alarms times theirs smallest? | It needs the prices, which only the business knows | Two prices | The decision itself |
| recall@k, precision@k, MRR, nDCG@k | Did we find the relevant documents, how much noise came with them, how high is the first hit, are the best ones on top? | Anything the golden set does not cover; nDCG needs graded relevance | A golden set: queries with the ids of the documents that answer them | Search and RAG retrieval |
| BLEU, ROUGE-L | How much wording does the answer share with a reference? | Truth: a correct paraphrase scores near zero | One or more reference answers | Machine translation, summarisation, legacy pipelines |
| Embedding similarity (BERTScore) | How close in meaning is the answer to a reference? | It still needs a reference, and closeness is not correctness | A reference and an embedding model | Generation with references |
| LLM judge, calibrated with Cohen's kappa | Does a rubric-following model agree with human graders beyond chance? | The judge's own biases (position, verbosity, self-enhancement) and chance agreement | A human-labelled sample and a rubric | Open-ended generation and agent evaluation |
How to choose
Name the question first, then the metric that answers it.
- A classifier with a threshold in production: report precision and recall at that threshold, and pick the threshold by pricing the two errors. In this lesson, pricing a miss at 10 and a false alarm at 1 drops the threshold and lifts recall from 0.30 to 0.88; the reverse prices raise it until precision is 1.00. Never lead with accuracy on imbalanced data.
- Comparing models before any threshold exists: ROC-AUC. If positives are rare, look at the precision-recall curve as well.
- Retrieval, including the retrieval half of RAG: recall@k with k set to the number of chunks you actually pass to the model, because the model cannot use what was not retrieved. Add MRR or nDCG when position matters.
- Generation: a code-based check wherever the answer can be verified (a test passes, a number matches, the JSON parses). Overlap metrics only where the task is nearly verbatim, such as translation. Everything else, an LLM judge with a written rubric, calibrated against people with kappa before it grades anything at scale.
- Whatever you pick, a metric is meaningless without its conditions: the threshold, the k, the golden set, the sample the judge was checked on. Report them next to the number, every time.
What it costs
Metrics cost labels, not compute. Classification needs labelled examples and, for the threshold, two prices you must extract from whoever owns the consequences. Retrieval needs a golden set of real queries paired with the documents that answer them, which is hours of a knowledgeable person's time and the single best investment in a RAG system, because you can then rerun recall@k after every change to chunking, the embedding model or the reranker. Generation costs the most: reference answers for overlap metrics, or human labels for the sample a judge is calibrated on, plus a model call per graded output for the judge itself, and the calibration is repeated every time the judge model or the rubric changes. The cost of the wrong metric is the one that matters: it is the gap between the dashboard's 94% and the four frauds in ten that walked through.
What breaks
- Accuracy on imbalanced data. 99% for a model that does nothing. Use precision and recall.
- A number with no threshold. "What's the accuracy?" is incomplete without "at what threshold?"; every classification metric changes when the cut-off moves.
- AUC on rare events. This lesson's classifier has an AUC of 0.89 and, at 80% recall, only about a third of its flags are real against a 10% base rate. Read the precision-recall curve.
- F1 when the errors cost differently. The harmonic mean pulls toward the smaller of the two: precision 1.0 with recall 0.1 gives F1 0.18. If a miss costs ten times a false alarm, F1 is the wrong target; price the errors instead.
- recall@k at the wrong k. Recall@20 is no comfort when you pass five chunks to the model. Measure at the k you serve.
- Overlap metrics on paraphrase. The Monday answer, factually wrong, scores 0.91 on ROUGE-L. Overlap measures wording, not truth.
- Raw agreement for a judge. 90% agreement and kappa 0 describe the same lazy judge. Kappa above about 0.6 is usually read as substantial agreement; below it, fix the rubric, the examples or the judge model.
- A judge that drifts. A new judge model or an edited rubric is a new judge. Re-run the calibration.
In the wild
scikit-learn ships the classification family as
precision_score, recall_score, f1_score, roc_auc_score,
average_precision_score and cohen_kappa_score, and its metrics guide is
in Further reading. In retrieval, the BEIR benchmark (Thakur et al., 2021)
compares lexical, sparse, dense, late-interaction and reranking systems
across 18 datasets. BLEU (Papineni et al., 2002) and ROUGE (Lin, 2004) are
still the reported numbers in translation and summarisation, with BERTScore
(Zhang et al., 2019) as the embedding-based successor. For judging, Zheng
et al. (2023) found that strong LLM judges reach over 80% agreement with
humans on MT-Bench, the level humans reach with each other, and named the
position, verbosity and self-enhancement biases every judge pipeline now
guards against. primer.agents.evals turns these metrics into a release
gate for an agent.
Go deeper
Level 2 builds each family from a hand-sized example: the four cells of the confusion matrix and the accuracy trap, the ROC curve as a walk whose area counts correctly ordered pairs, the cheapest threshold under two prices, the four retrieval metrics on one five-document list, BLEU and ROUGE-L on the Monday sentence, and Cohen's kappa on twenty essays. If you only needed to pick the metric and read it honestly, you are done.
Level 2
How it works, from scratch
A loss function is what the model optimizes during training; a metric is what you use to decide whether it's working. Picking the wrong metric is one of the most common ways a project looks great offline and fails in production. This lesson builds every metric from scratch in three families:
- Classification: precision, recall, F1, accuracy, the confusion matrix, ROC-AUC, precision-recall curves, and choosing a threshold by the cost of each kind of error.
- Retrieval: recall@k, precision@k, MRR and nDCG, the metrics for search and retrieval-augmented generation (RAG).
- Generation: BLEU and ROUGE-L (word overlap), why they fail on language-model output, and how to check an automated judge against people with Cohen's kappa.
Chapter 1
Classification: precision, recall and the confusion matrix
Everyday picture You're fishing for trout in a lake that also holds old boots. Precision asks: of everything in your net, how much is trout? Recall asks: of all the trout in the lake, how many did you catch? A tiny net held in one good spot has high precision and low recall; dragging a huge net across the whole lake has high recall and a lot of boots.
Tiny worked example: fraud detection. Of 100 transactions, 10 are fraud. The model flags 8, and 6 of those are truly fraud. So 6 hits (true positives, TP), 2 false alarms (false positives, FP), 4 misses (false negatives, FN) and 88 correct passes (true negatives, TN).
| Metric | Calculation | Result |
|---|---|---|
| Precision | 6 correct ÷ 8 flagged | 75% |
| Recall | 6 found ÷ 10 actual fraud | 60% |
| F1 | 2 × 0.75 × 0.60 ÷ (0.75 + 0.60) | 67% |
| Accuracy | (6 + 88 correct) ÷ 100 | 94% |
Accuracy looks great at 94% while the model misses 4 in 10 frauds.
Figure 1 · Diagram
flowchart LR
X[Item] --> M[Model]
M --> S[Score<br/>e.g. 0.73]
S --> T{Score >= threshold?}
T -->|yes| P[Flagged positive]
T -->|no| N[Not flagged]
P --> CM[Confusion matrix<br/>TP / FP / FN / TN]
N --> CM
L[True label] --> CM
CM --> MET[Precision, recall,<br/>F1, accuracy]
| predicted positive | predicted negative | |
|---|---|---|
| actual positive | TP = 6 | FN = 4 (missed) |
| actual negative | FP = 2 (false alarm) | TN = 88 |
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| TP, FP, FN, TN | counts of hits, false alarms, misses and correct passes | whole numbers |
| , | precision and recall | 0 … 1 |
| the harmonic mean of P and R: an average that is dragged towards the smaller of the two | 0 … 1 | |
| total number of items, TP + FP + FN + TN | whole number | |
| fraction bar | "divided by" |
In words: precision is hits over everything flagged; recall is hits over everything that should have been flagged; F1 is twice their product over their sum; accuracy is everything right over everything.
On the worked example: P = 6/(6 + 2) = 0.75; R = 6/(6 + 4) = 0.60; F1 = 2 × 0.75 × 0.60 / 1.35 = 0.667; accuracy = (6 + 88)/100 = 0.94. The harmonic mean punishes imbalance: P = 1.0 with R = 0.1 gives F1 = 0.18, not the ordinary average of 0.55.
Level 3: in Python
TP, FP, FN, TN = 6, 2, 4, 88
N = TP + FP + FN + TN
# precision, recall
P, R = TP / (TP + FP), TP / (TP + FN)
P, R # → (0.75, 0.6)
# F1, accuracy
round(2 * P * R / (P + R), 3), (TP + TN) / N # → (0.667, 0.94)
# F1 when P = 1.0 and R = 0.1
round(2 * 1.0 * 0.1 / (1.0 + 0.1), 2) # → 0.18
In code: confusion sorts labels and decisions into a Confusion,
whose Confusion.precision, Confusion.recall, Confusion.f1 and
Confusion.accuracy are the four formulas. fraud_example rebuilds the
100 transactions above, and flag_nothing_trap builds the model below that
never flags anything.
Why it matters in practice: the accuracy trap. If 1% of transactions are fraud, a model that flags nothing is 99% accurate and completely useless (recall 0). On imbalanced data, never lead with accuracy.
Chapter 2
ROC-AUC: how well does the model rank?
Everyday picture A smoke alarm has a sensitivity dial. Turn it up and it catches every fire but also shrieks at toast; turn it down and it stays quiet but might miss a real fire. The ROC curve draws every dial setting at once, and AUC scores the alarm across all of them.
Tiny worked example Two frauds scored 0.9 and 0.4, two legitimate transactions scored 0.6 and 0.1. Compare every (fraud, legitimate) pair: 0.9 > 0.6 ✓, 0.9 > 0.1 ✓, 0.4 > 0.6 ✗, 0.4 > 0.1 ✓. Three of four pairs are ranked correctly, so AUC = 0.75.
Figure 4 · Diagram
flowchart TD
A[Sort items by score, highest first] --> B[Start with threshold above every score<br/>nothing flagged: point 0,0]
B --> C[Lower the threshold past the next distinct score]
C --> D{Which items became flagged?}
D -->|a positive| E[Step up: TPR rises]
D -->|a negative| F[Step right: FPR rises]
D -->|a tie of both| G[Diagonal step: half credit]
E --> H{Everything flagged?}
F --> H
G --> H
H -->|no| C
H -->|yes| I[End at point 1,1<br/>AUC = area under the path]
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| TPR | true-positive rate: recall, the share of positives flagged | 0 … 1 |
| FPR | false-positive rate: the share of negatives wrongly flagged | 0 … 1 |
| , (in the AUC sum) | the set of positive items and the set of negative items | sets |
| , | how many items each set holds | whole numbers |
| , | the model's score for positive p and negative n | real |
| 1 if the statement inside is true, else 0 | 0 or 1 | |
| add over every (positive, negative) pair |
In words: AUC is the share of (positive, negative) pairs in which the positive gets the higher score, counting ties as half.
On the worked example: 2 positives × 2 negatives = 4 pairs; 3 are ordered correctly; AUC = 3/4 = 0.75. One point on the curve: at threshold 0.5 the model flags the 0.9 fraud and the 0.6 legitimate transaction, so TPR = 1/(1 + 1) = 0.5 and FPR = 1/(1 + 1) = 0.5.
Level 3: in Python
# scores of the positives (frauds)
s_P = [0.9, 0.4]
# scores of the negatives (legitimate)
s_N = [0.6, 0.1]
t = 0.5
TP, FN = sum(s >= t for s in s_P), sum(s < t for s in s_P)
FP, TN = sum(s >= t for s in s_N), sum(s < t for s in s_N)
# TPR, FPR
TP / (TP + FN), FP / (FP + TN) # → (0.5, 0.5)
pairs = sum((s_p > s_n) + 0.5 * (s_p == s_n) for s_p in s_P for s_n in s_N)
# AUC: the share of pairs ranked correctly
pairs / (len(s_P) * len(s_N)) # → 0.75
Figure 2 · Drawn from the lesson's code
ROC curve bowing well above the coin-flip diagonal: the classifier ranks a random positive above a random negative 89% of the time (AUC 0.89)
synthetic_scores()
(positives shifted 1.5 standard deviations above negatives); the dashed
diagonal is a coin flip. The shaded area is the AUC, about 0.89: pick one
fraud and one legitimate transaction at random and the model scores the
fraud higher 89% of the time. Where the curve bends is where a sensible
threshold lives.Figure 3 · Drawn from the lesson's code
Precision-recall curve falling as recall rises: at 80% recall only about a third of the flags are real, against a 10% base rate
In code: roc_curve takes the walk, roc_auc measures the area under
it with auc_trapezoid, and roc_auc_rank counts correctly ordered pairs
instead; Confusion.fpr is FPR, and pr_curve gives the precision-recall
points.
Why it matters in practice. AUC compares models independently of any threshold. On heavily imbalanced data, look at the precision-recall curve too, because FPR stays tiny even when false alarms swamp the true positives.
Chapter 3
Picking the threshold: price your errors
Everyday picture A hospital screening test and a spam filter want opposite things. Missing a disease is terrible, so the screen flags generously; losing an important email is annoying, so the spam filter flags cautiously. Same maths, different prices.
Tiny worked example With the four scores above (frauds 0.9 and 0.4, legitimate 0.6 and 0.1): if a miss costs 10 and a false alarm costs 1, the cheapest rule is "flag at 0.4 or above", catching both frauds and wrongly flagging one legitimate transaction (total cost 1). If a false alarm costs 10 and a miss 1, the cheapest rule is "flag only 0.9", missing one fraud but raising no false alarms (total cost 1).
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| the threshold: flag items scoring at least t | real | |
| , | misses and false alarms at that threshold | whole numbers |
| , | the price of one miss and of one false alarm | ≥ 0 |
| "the value of t that makes this smallest" | ||
| the best threshold | real |
In words: the total cost at a threshold is misses times their price plus false alarms times theirs; pick the threshold where that total is lowest.
On the worked example: c_FN = 10, c_FP = 1: at t = 0.4, FN = 0 and FP = 1, cost 1, the minimum. The other thresholds cost 10 (t = 0.9, one miss), 11 (t = 0.6, one miss and one false alarm) and 2 (t = 0.1, two false alarms).
Level 3: in Python
frauds, legit = [0.9, 0.4], [0.6, 0.1]
c_FN, c_FP = 10, 1
def cost(t):
# frauds below t are missed
FN = sum(s < t for s in frauds)
# legitimate ones at or above t are false alarms
FP = sum(s >= t for s in legit)
return c_FN * FN + c_FP * FP
[cost(t) for t in (0.9, 0.6, 0.4, 0.1)] # → [10, 11, 1, 2]
# t* = argmin_t cost(t)
min((0.9, 0.6, 0.4, 0.1), key=cost) # → 0.4
Figure 5 · Drawn from the lesson's code
Total cost against threshold for three error prices: the cheapest threshold moves left when misses cost more and right when false alarms cost more
In code: best_threshold_by_cost tries every distinct score as t
(plus "flag nothing") and keeps the cheapest.
Chapter 4
Retrieval metrics: judging a search result list
Everyday picture You ask a librarian for books on a topic and get a stack of five. Did the stack include the books that matter (recall)? How much of the stack is useful (precision)? Is a good book on top, or do you dig for it (reciprocal rank)? Are the best books nearest the top (nDCG)?
Tiny worked example One query. The system returns d7, d3, d9, d1, d4 in that order. The relevant documents are d3 (very relevant, grade 3), d4 (grade 2) and d8 (grade 1, never returned).
- recall@5 = 2 found of 3 relevant = 0.667
- precision@5 = 2 relevant of 5 returned = 0.4
- reciprocal rank = first relevant result at rank 2, so 1/2 = 0.5
- nDCG@5 = 0.560 (worked below)
Figure 7 · Diagram
flowchart LR G[Golden set<br/>query + relevant doc ids] --> R[Retriever] R --> K[Ranked top-k list] K --> RK["recall@k<br/>did we find them?"] K --> PK["precision@k<br/>how much noise?"] K --> RR[MRR<br/>how high is the first hit?] K --> ND["nDCG@k<br/>are the best ones on top?"] G --> RK & PK & RR & ND
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| how many top results you look at | e.g. 5 | |
| "in both": documents that are relevant and in the top k | set | |
| how many items a set holds | ||
| , | the set of test queries, and one query | |
| position (1 = top) of the first relevant result for query q | 1, 2, … | |
| the relevance grade of the result at position i (0 if irrelevant) | e.g. 0 … 3 | |
| logarithm base 2: how many times you halve i + 1 to reach 1. It grows slowly, so it is a gentle position discount | 1 at rank 1, 1.58 at rank 2 | |
| IDCG | "ideal DCG": the DCG of the same grades sorted best-first, so nDCG tops out at 1 |
In words: recall@k is the share of relevant documents found in the top k; MRR averages one-over-the-rank of the first hit; DCG adds each result's grade, discounted by the log of its position; nDCG divides by the best possible DCG.
On the worked example: recall@5 = 2/3 = 0.667; with one query whose first hit is at rank 2, MRR = 1/2 = 0.5. DCG = 3/log₂(3) + 2/log₂(6) = 1.8928 + 0.7737 = 2.6665 (d3 at rank 2, d4 at rank 5). Ideal order d3, d4, d8: IDCG = 3/1 + 2/1.585 + 1/2 = 4.7619. nDCG = 2.6665/4.7619 = 0.560.
Level 3: in Python
import math
ranked = ["d7", "d3", "d9", "d1", "d4"]
# relevance grades; missing means 0
rel = {"d3": 3, "d4": 2, "d8": 1}
k = 5
# recall@k
round(len(set(rel) & set(ranked[:k])) / len(rel), 3) # → 0.667
rank_q = next(i for i, doc in enumerate(ranked, start=1) if doc in rel)
# MRR over a single query
1 / rank_q # → 0.5
DCG = sum(rel.get(doc, 0) / math.log2(i + 1) for i, doc in enumerate(ranked[:k], start=1))
# the same grades, best first
best = sorted(rel.values(), reverse=True)[:k]
IDCG = sum(g / math.log2(i + 1) for i, g in enumerate(best, start=1))
print(f"{DCG:.4f} {IDCG:.4f} {DCG / IDCG:.3f}") # → 2.6665 4.7619 0.560
Figure 6 · Drawn from the lesson's code
DCG weight by rank: rank 1 counts fully, rank 2 counts 63% and rank 10 still counts 29%
In code: recall_at_k and precision_at_k score one ranked list,
reciprocal_rank finds its first hit and mean_reciprocal_rank averages
that over queries; dcg_at_k adds the discounted grades and ndcg_at_k
divides by the ideal ordering's DCG.
Why it matters in practice. For RAG, recall@k (with k = the number of chunks you pass to the model) usually matters most: the model can't use a passage that wasn't retrieved, and no prompt fixes that.
Chapter 5
Generation metrics: why word overlap misleads
Everyday picture Grading essays by counting how many words each one shares with the answer key. A student who copies the key's phrasing but gets the key fact wrong scores well; a student who explains it correctly in their own words scores badly.
Tiny worked example Reference: "The meeting was moved to Friday because the manager is sick."
| candidate | BLEU | ROUGE-L |
|---|---|---|
| "Since the boss is ill, the team rescheduled the meeting for Friday." (correct) | 0.16 | 0.26 |
| "The meeting was moved to Monday because the manager is sick." (wrong) | 0.73 | 0.91 |
BLEU (from machine translation) multiplies together how many of the candidate's 1-, 2-, 3- and 4-word sequences appear in the reference, with a penalty for being too short. ROUGE-L (from summarisation) measures the longest sequence of words the two share in the same order, gaps allowed.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| share of the candidate's n-word sequences also found in the reference, each counted at most as often as the reference has it | 0 … 1 | |
| , | natural log and its inverse; exp of the average log is the geometric mean of the four | |
| BP | brevity penalty: 1 if the candidate is longer than the reference, else e^(1 − ref length / candidate length) | 0 … 1 |
| LCS | longest common subsequence: the longest run of words appearing in both texts in the same order | |
| , | LCS length over candidate length, and over reference length | 0 … 1 |
In words: BLEU is the geometric mean of the n-gram match rates, scaled down for short answers; ROUGE-L is the F1 of the longest shared in-order word sequence.
On a hand example: candidate "a b c d", reference "a c d e": the LCS is "a c d", 3 words, so P = 3/4, R = 3/4, ROUGE-L = 0.75. And "the the the the" against "the cat is here" scores unigram precision 1/4: "the" is credited only as often as the reference contains it.
With the numbers: the Monday answer and the reference are both 11
words (full stop and capitals dropped), so BP = 1. It matches 10 of its 11
words, 8 of its 10 word pairs, 6 of its 9 triples and 4 of its 8 four-word
runs. bleu adds 1 to the top and bottom for n ≥ 2 (smoothing, so one
missing four-word run can't zero the score), giving p = 10/11, 9/11, 7/10,
5/9, and BLEU = exp(¼(ln 0.909 + ln 0.818 + ln 0.7 + ln 0.556)) = 0.73.
Its longest shared in-order run is 10 words, so ROUGE-L = 2 · (10/11) ·
(10/11) / (20/11) = 0.91.
In Python:
import math
ref = "the meeting was moved to friday because the manager is sick".split()
cand = "the meeting was moved to monday because the manager is sick".split()
# every run of n words, in order
def grams(words, n):
return [tuple(words[i:i + n]) for i in range(len(words) - n + 1)]
def p(n):
c, r = grams(cand, n), grams(ref, n)
# clipped: at most as often as ref has it
hits = sum(min(c.count(g), r.count(g)) for g in set(c))
# add-one smoothing for n ≥ 2
s = 1 if n > 1 else 0
return (hits + s) / (len(c) + s)
[round(p(n), 3) for n in range(1, 5)] # → [0.909, 0.818, 0.7, 0.556]
BP = 1.0 if len(cand) > len(ref) else math.exp(1 - len(ref) / len(cand))
# BLEU
round(BP * math.exp(sum(math.log(p(n)) for n in range(1, 5)) / 4), 2) # → 0.73
# every word but "monday", in order
P_LCS, R_LCS = 10 / len(cand), 10 / len(ref)
# ROUGE-L
round(2 * P_LCS * R_LCS / (P_LCS + R_LCS), 2) # → 0.91
# the hand example: LCS "a c d"
P_LCS, R_LCS = 3 / 4, 3 / 4
2 * P_LCS * R_LCS / (P_LCS + R_LCS) # → 0.75
Figure 8 · Drawn from the lesson's code
BLEU and ROUGE-L bars: the correct paraphrase scores lowest while the wrong answer that copies the wording scores almost as high as an exact copy
In code: bleu clips each n-gram count, takes the geometric mean and
applies the brevity penalty; rouge_l finds the longest common subsequence
and returns its F1.
Why it matters in practice. Teams grade open-ended output with an LLM as a judge, a model following an explicit rubric, and embedding-based scores like BERTScore do somewhat better than overlap. But a judge must earn trust first (next section).
Chapter 6
Calibrating a judge: Cohen's kappa
Everyday picture Two teachers grade the same 20 essays pass/fail and agree on 18. Impressive? Not if 18 of the essays were obvious passes: a careful teacher who reads every essay (and passes those 18) and a lazy one who stamps "pass" on everything without reading would also agree on 18. Cohen's kappa subtracts the agreement you'd expect from luck.
Tiny worked example Human labels: 18 pass, 2 fail. A lazy judge says "pass" to everything. Raw agreement is 90%, and chance agreement (both say pass at their own rates) is also 0.9 × 1.0 + 0.1 × 0 = 90%, so kappa = 0. A second example: labels (y, y, n, n) vs. (y, n, n, n) agree 3/4 of the time; chance predicts 0.5 × 0.25 + 0.5 × 0.75 = 0.5; kappa = 0.5.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| kappa: agreement beyond chance | ≤ 1; 0 = chance, 1 = perfect | |
| observed agreement: share of items both raters labelled the same | 0 … 1 | |
| agreement expected by chance | 0 … 1 | |
| one possible label (pass, fail…) | ||
| , | how often rater A (and rater B) uses label ℓ | 0 … 1 |
In words: kappa is how far agreement beats chance, as a share of the most it could beat chance by.
On the worked example: p_o = 0.75, p_e = 0.5, κ = (0.75 − 0.5)/(1 − 0.5) = 0.5.
Level 3: in Python
# rater A's labels
A = ["y", "y", "n", "n"]
# rater B's labels
B = ["y", "n", "n", "n"]
p_o = sum(a == b for a, b in zip(A, B)) / len(A)
# Σ_ℓ p_A(ℓ) p_B(ℓ)
p_e = sum((A.count(ell) / len(A)) * (B.count(ell) / len(B)) for ell in ("y", "n"))
p_o, p_e # → (0.75, 0.5)
# κ
(p_o - p_e) / (1 - p_e) # → 0.5
Figure 9 · Diagram
flowchart LR S[Sample of real outputs] --> H[Humans label<br/>pass / fail] S --> J[LLM judge + rubric<br/>labels pass / fail] H --> K[Cohen's kappa] J --> K K -->|kappa high enough| U[Use the judge at scale] K -->|too low| F[Fix rubric, examples<br/>or judge model] F --> J
In code: cohens_kappa computes p_o and p_e from two lists of labels
and returns κ.
Why it matters in practice. Kappa above about 0.6 is usually considered substantial agreement. Re-check it whenever the judge model or rubric changes.
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: A fraud model has 94% accuracy. Is it good?Think it through, then reveal
A: You can't tell from accuracy. If fraud is 10% of traffic, check recall and precision. In the worked example recall is only 60%, so 4 in 10 frauds slip through.
Question 2Q: When do you optimize for precision vs. recall?Think it through, then reveal
A: By the cost of each error. If a miss is expensive (fraud, cancer screening, a relevant legal document), favor recall. If a false alarm is expensive (blocking a good customer, paging an engineer at 3 a.m.), favor precision. Set the threshold to minimize expected cost.
Question 3Q: What does an AUC of 0.8 mean?Think it through, then reveal
A: A randomly chosen positive gets a higher score than a randomly chosen negative 80% of the time. It measures ranking quality across all thresholds; 0.5 is random.
Question 4Q: Which retrieval metric matters most for RAG, and why?Think it through, then reveal
A: Recall@k, where k is the number of chunks you pass to the model. If the right passage isn't in the top k, no prompt engineering can recover the answer. Add MRR or nDCG when position within the context matters.
Question 5Q: Why not grade LLM answers with BLEU or ROUGE?Think it through, then reveal
A: They measure surface overlap with one reference. A correct paraphrase scores low and a wrong answer that reuses the reference's words scores high. Use code-based checks where possible and a rubric-driven LLM judge otherwise, validated against human labels.
Question 6Q: Your LLM judge agrees with humans 90% of the time. Good enough?Think it through, then reveal
A: Not necessarily. If 90% of answers are "pass", a judge that always says "pass" also agrees 90%. Compute Cohen's kappa to correct for chance agreement.
Primary sources
The papers behind this lesson
Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023): Measured how well strong language models agree with human preferences when used as judges, and catalogued their biases (position, verbosity, self-preference).
Read the annotated companion →The paper ↗Papineni et al., BLEU: a Method for Automatic Evaluation of Machine Translation (2002): Introduced clipped n-gram precision with a brevity penalty.
The paper ↗Lin, ROUGE: A Package for Automatic Evaluation of Summaries (2004): Introduced recall-oriented overlap measures, including the LCS-based ROUGE-L.
The paper ↗Zhang et al., BERTScore: Evaluating Text Generation with BERT (2019): Compared candidate and reference by embedding similarity rather than exact word overlap.
The paper ↗Researcher's shelf
Further reading
- scikit-learn, Metrics and scoring: https://scikit-learn.org/stable/modules/model_evaluation.html
- Receiver operating characteristic (Wikipedia): https://en.wikipedia.org/wiki/Receiver_operating_characteristic
- Discounted cumulative gain (Wikipedia): https://en.wikipedia.org/wiki/Discounted_cumulative_gain
- Cohen's kappa (Wikipedia): https://en.wikipedia.org/wiki/Cohen%27s_kappa
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.