At a glance
Key takeaways
- Precision: of what I flagged, how much was right. Recall: of what mattered, how much I found. F1 balances them.
- Accuracy lies on imbalanced data. Flagging nothing on 1% fraud scores 99%.
- ROC-AUC is the probability a random positive outranks a random negative.
- Choose the threshold by what each kind of error costs.
- For RAG, measure recall@k first. What isn't retrieved can't be used.
- BLEU/ROUGE punish correct paraphrases; use an LLM judge and calibrate it against humans with Cohen's kappa.
Level 2
How it works, from scratch
A loss function is what the model optimizes during training; a metric is what you use to decide whether it's working. Picking the wrong metric is one of the most common ways a project looks great offline and fails in production. This lesson builds every metric from scratch in three families:
- Classification: precision, recall, F1, accuracy, the confusion matrix, ROC-AUC, precision-recall curves, and choosing a threshold by the cost of each kind of error.
- Retrieval: recall@k, precision@k, MRR and nDCG, the metrics for search and retrieval-augmented generation (RAG).
- Generation: BLEU and ROUGE-L (word overlap), why they fail on language-model output, and how to check an automated judge against people with Cohen's kappa.
Chapter 1
Classification: precision, recall and the confusion matrix
Everyday picture You're fishing for trout in a lake that also holds old boots. Precision asks: of everything in your net, how much is trout? Recall asks: of all the trout in the lake, how many did you catch? A tiny net held in one good spot has high precision and low recall; dragging a huge net across the whole lake has high recall and a lot of boots.
Tiny worked example: fraud detection. Of 100 transactions, 10 are fraud. The model flags 8, and 6 of those are truly fraud. So 6 hits (true positives, TP), 2 false alarms (false positives, FP), 4 misses (false negatives, FN) and 88 correct passes (true negatives, TN).
| Metric | Calculation | Result |
|---|---|---|
| Precision | 6 correct ÷ 8 flagged | 75% |
| Recall | 6 found ÷ 10 actual fraud | 60% |
| F1 | 2 × 0.75 × 0.60 ÷ (0.75 + 0.60) | 67% |
| Accuracy | (6 + 88 correct) ÷ 100 | 94% |
Accuracy looks great at 94% while the model misses 4 in 10 frauds.
Figure 1 · Diagram
flowchart LR
X[Item] --> M[Model]
M --> S[Score<br/>e.g. 0.73]
S --> T{Score >= threshold?}
T -->|yes| P[Flagged positive]
T -->|no| N[Not flagged]
P --> CM[Confusion matrix<br/>TP / FP / FN / TN]
N --> CM
L[True label] --> CM
CM --> MET[Precision, recall,<br/>F1, accuracy]
| predicted positive | predicted negative | |
|---|---|---|
| actual positive | TP = 6 | FN = 4 (missed) |
| actual negative | FP = 2 (false alarm) | TN = 88 |
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| TP, FP, FN, TN | counts of hits, false alarms, misses and correct passes | whole numbers |
| , | precision and recall | 0 … 1 |
| the harmonic mean of P and R: an average that is dragged towards the smaller of the two | 0 … 1 | |
| total number of items, TP + FP + FN + TN | whole number | |
| fraction bar | "divided by" |
In words: precision is hits over everything flagged; recall is hits over everything that should have been flagged; F1 is twice their product over their sum; accuracy is everything right over everything.
On the worked example: P = 6/(6 + 2) = 0.75; R = 6/(6 + 4) = 0.60; F1 = 2 × 0.75 × 0.60 / 1.35 = 0.667; accuracy = (6 + 88)/100 = 0.94. The harmonic mean punishes imbalance: P = 1.0 with R = 0.1 gives F1 = 0.18, not the ordinary average of 0.55.
Level 3: in Python
TP, FP, FN, TN = 6, 2, 4, 88
N = TP + FP + FN + TN
# precision, recall
P, R = TP / (TP + FP), TP / (TP + FN)
P, R # → (0.75, 0.6)
# F1, accuracy
round(2 * P * R / (P + R), 3), (TP + TN) / N # → (0.667, 0.94)
# F1 when P = 1.0 and R = 0.1
round(2 * 1.0 * 0.1 / (1.0 + 0.1), 2) # → 0.18
In code: confusion sorts labels and decisions into a Confusion,
whose Confusion.precision, Confusion.recall, Confusion.f1 and
Confusion.accuracy are the four formulas. fraud_example rebuilds the
100 transactions above, and flag_nothing_trap builds the model below that
never flags anything.
Why it matters in practice: the accuracy trap. If 1% of transactions are fraud, a model that flags nothing is 99% accurate and completely useless (recall 0). On imbalanced data, never lead with accuracy.
Chapter 2
ROC-AUC: how well does the model rank?
Everyday picture A smoke alarm has a sensitivity dial. Turn it up and it catches every fire but also shrieks at toast; turn it down and it stays quiet but might miss a real fire. The ROC curve draws every dial setting at once, and AUC scores the alarm across all of them.
Tiny worked example Two frauds scored 0.9 and 0.4, two legitimate transactions scored 0.6 and 0.1. Compare every (fraud, legitimate) pair: 0.9 > 0.6 ✓, 0.9 > 0.1 ✓, 0.4 > 0.6 ✗, 0.4 > 0.1 ✓. Three of four pairs are ranked correctly, so AUC = 0.75.
Figure 2 · Diagram
flowchart TD
A[Sort items by score, highest first] --> B[Start with threshold above every score<br/>nothing flagged: point 0,0]
B --> C[Lower the threshold past the next distinct score]
C --> D{Which items became flagged?}
D -->|a positive| E[Step up: TPR rises]
D -->|a negative| F[Step right: FPR rises]
D -->|a tie of both| G[Diagonal step: half credit]
E --> H{Everything flagged?}
F --> H
G --> H
H -->|no| C
H -->|yes| I[End at point 1,1<br/>AUC = area under the path]
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| TPR | true-positive rate: recall, the share of positives flagged | 0 … 1 |
| FPR | false-positive rate: the share of negatives wrongly flagged | 0 … 1 |
| , (in the AUC sum) | the set of positive items and the set of negative items | sets |
| , | how many items each set holds | whole numbers |
| , | the model's score for positive p and negative n | real |
| 1 if the statement inside is true, else 0 | 0 or 1 | |
| add over every (positive, negative) pair |
In words: AUC is the share of (positive, negative) pairs in which the positive gets the higher score, counting ties as half.
On the worked example: 2 positives × 2 negatives = 4 pairs; 3 are ordered correctly; AUC = 3/4 = 0.75. One point on the curve: at threshold 0.5 the model flags the 0.9 fraud and the 0.6 legitimate transaction, so TPR = 1/(1 + 1) = 0.5 and FPR = 1/(1 + 1) = 0.5.
Level 3: in Python
# scores of the positives (frauds)
s_P = [0.9, 0.4]
# scores of the negatives (legitimate)
s_N = [0.6, 0.1]
t = 0.5
TP, FN = sum(s >= t for s in s_P), sum(s < t for s in s_P)
FP, TN = sum(s >= t for s in s_N), sum(s < t for s in s_N)
# TPR, FPR
TP / (TP + FN), FP / (FP + TN) # → (0.5, 0.5)
pairs = sum((s_p > s_n) + 0.5 * (s_p == s_n) for s_p in s_P for s_n in s_N)
# AUC: the share of pairs ranked correctly
pairs / (len(s_P) * len(s_N)) # → 0.75
Figure 3 · Drawn from the lesson's code
ROC curve bowing well above the coin-flip diagonal: the classifier ranks a random positive above a random negative 89% of the time (AUC 0.89)
synthetic_scores()
(positives shifted 1.5 standard deviations above negatives); the dashed
diagonal is a coin flip. The shaded area is the AUC, about 0.89: pick one
fraud and one legitimate transaction at random and the model scores the
fraud higher 89% of the time. Where the curve bends is where a sensible
threshold lives.Figure 4 · Drawn from the lesson's code
Precision-recall curve falling as recall rises: at 80% recall only about a third of the flags are real, against a 10% base rate
In code: roc_curve takes the walk, roc_auc measures the area under
it with auc_trapezoid, and roc_auc_rank counts correctly ordered pairs
instead; Confusion.fpr is FPR, and pr_curve gives the precision-recall
points.
Why it matters in practice. AUC compares models independently of any threshold. On heavily imbalanced data, look at the precision-recall curve too, because FPR stays tiny even when false alarms swamp the true positives.
Chapter 3
Picking the threshold: price your errors
Everyday picture A hospital screening test and a spam filter want opposite things. Missing a disease is terrible, so the screen flags generously; losing an important email is annoying, so the spam filter flags cautiously. Same maths, different prices.
Tiny worked example With the four scores above (frauds 0.9 and 0.4, legitimate 0.6 and 0.1): if a miss costs 10 and a false alarm costs 1, the cheapest rule is "flag at 0.4 or above", catching both frauds and wrongly flagging one legitimate transaction (total cost 1). If a false alarm costs 10 and a miss 1, the cheapest rule is "flag only 0.9", missing one fraud but raising no false alarms (total cost 1).
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| the threshold: flag items scoring at least t | real | |
| , | misses and false alarms at that threshold | whole numbers |
| , | the price of one miss and of one false alarm | ≥ 0 |
| "the value of t that makes this smallest" | ||
| the best threshold | real |
In words: the total cost at a threshold is misses times their price plus false alarms times theirs; pick the threshold where that total is lowest.
On the worked example: c_FN = 10, c_FP = 1: at t = 0.4, FN = 0 and FP = 1, cost 1, the minimum. The other thresholds cost 10 (t = 0.9, one miss), 11 (t = 0.6, one miss and one false alarm) and 2 (t = 0.1, two false alarms).
Level 3: in Python
frauds, legit = [0.9, 0.4], [0.6, 0.1]
c_FN, c_FP = 10, 1
def cost(t):
# frauds below t are missed
FN = sum(s < t for s in frauds)
# legitimate ones at or above t are false alarms
FP = sum(s >= t for s in legit)
return c_FN * FN + c_FP * FP
[cost(t) for t in (0.9, 0.6, 0.4, 0.1)] # → [10, 11, 1, 2]
# t* = argmin_t cost(t)
min((0.9, 0.6, 0.4, 0.1), key=cost) # → 0.4
Figure 5 · Drawn from the lesson's code
Total cost against threshold for three error prices: the cheapest threshold moves left when misses cost more and right when false alarms cost more
In code: best_threshold_by_cost tries every distinct score as t
(plus "flag nothing") and keeps the cheapest.
Chapter 4
Retrieval metrics: judging a search result list
Everyday picture You ask a librarian for books on a topic and get a stack of five. Did the stack include the books that matter (recall)? How much of the stack is useful (precision)? Is a good book on top, or do you dig for it (reciprocal rank)? Are the best books nearest the top (nDCG)?
Tiny worked example One query. The system returns d7, d3, d9, d1, d4 in that order. The relevant documents are d3 (very relevant, grade 3), d4 (grade 2) and d8 (grade 1, never returned).
- recall@5 = 2 found of 3 relevant = 0.667
- precision@5 = 2 relevant of 5 returned = 0.4
- reciprocal rank = first relevant result at rank 2, so 1/2 = 0.5
- nDCG@5 = 0.560 (worked below)
Figure 6 · Diagram
flowchart LR G[Golden set<br/>query + relevant doc ids] --> R[Retriever] R --> K[Ranked top-k list] K --> RK["recall@k<br/>did we find them?"] K --> PK["precision@k<br/>how much noise?"] K --> RR[MRR<br/>how high is the first hit?] K --> ND["nDCG@k<br/>are the best ones on top?"] G --> RK & PK & RR & ND
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| how many top results you look at | e.g. 5 | |
| "in both": documents that are relevant and in the top k | set | |
| how many items a set holds | ||
| , | the set of test queries, and one query | |
| position (1 = top) of the first relevant result for query q | 1, 2, … | |
| the relevance grade of the result at position i (0 if irrelevant) | e.g. 0 … 3 | |
| logarithm base 2: how many times you halve i + 1 to reach 1. It grows slowly, so it is a gentle position discount | 1 at rank 1, 1.58 at rank 2 | |
| IDCG | "ideal DCG": the DCG of the same grades sorted best-first, so nDCG tops out at 1 |
In words: recall@k is the share of relevant documents found in the top k; MRR averages one-over-the-rank of the first hit; DCG adds each result's grade, discounted by the log of its position; nDCG divides by the best possible DCG.
On the worked example: recall@5 = 2/3 = 0.667; with one query whose first hit is at rank 2, MRR = 1/2 = 0.5. DCG = 3/log₂(3) + 2/log₂(6) = 1.8928 + 0.7737 = 2.6665 (d3 at rank 2, d4 at rank 5). Ideal order d3, d4, d8: IDCG = 3/1 + 2/1.585 + 1/2 = 4.7619. nDCG = 2.6665/4.7619 = 0.560.
Level 3: in Python
import math
ranked = ["d7", "d3", "d9", "d1", "d4"]
# relevance grades; missing means 0
rel = {"d3": 3, "d4": 2, "d8": 1}
k = 5
# recall@k
round(len(set(rel) & set(ranked[:k])) / len(rel), 3) # → 0.667
rank_q = next(i for i, doc in enumerate(ranked, start=1) if doc in rel)
# MRR over a single query
1 / rank_q # → 0.5
DCG = sum(rel.get(doc, 0) / math.log2(i + 1) for i, doc in enumerate(ranked[:k], start=1))
# the same grades, best first
best = sorted(rel.values(), reverse=True)[:k]
IDCG = sum(g / math.log2(i + 1) for i, g in enumerate(best, start=1))
print(f"{DCG:.4f} {IDCG:.4f} {DCG / IDCG:.3f}") # → 2.6665 4.7619 0.560
Figure 7 · Drawn from the lesson's code
DCG weight by rank: rank 1 counts fully, rank 2 counts 63% and rank 10 still counts 29%
In code: recall_at_k and precision_at_k score one ranked list,
reciprocal_rank finds its first hit and mean_reciprocal_rank averages
that over queries; dcg_at_k adds the discounted grades and ndcg_at_k
divides by the ideal ordering's DCG.
Why it matters in practice. For RAG, recall@k (with k = the number of chunks you pass to the model) usually matters most: the model can't use a passage that wasn't retrieved, and no prompt fixes that.
Chapter 5
Generation metrics: why word overlap misleads
Everyday picture Grading essays by counting how many words each one shares with the answer key. A student who copies the key's phrasing but gets the key fact wrong scores well; a student who explains it correctly in their own words scores badly.
Tiny worked example Reference: "The meeting was moved to Friday because the manager is sick."
| candidate | BLEU | ROUGE-L |
|---|---|---|
| "Since the boss is ill, the team rescheduled the meeting for Friday." (correct) | 0.16 | 0.26 |
| "The meeting was moved to Monday because the manager is sick." (wrong) | 0.73 | 0.91 |
BLEU (from machine translation) multiplies together how many of the candidate's 1-, 2-, 3- and 4-word sequences appear in the reference, with a penalty for being too short. ROUGE-L (from summarisation) measures the longest sequence of words the two share in the same order, gaps allowed.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| share of the candidate's n-word sequences also found in the reference, each counted at most as often as the reference has it | 0 … 1 | |
| , | natural log and its inverse; exp of the average log is the geometric mean of the four | |
| BP | brevity penalty: 1 if the candidate is longer than the reference, else e^(1 − ref length / candidate length) | 0 … 1 |
| LCS | longest common subsequence: the longest run of words appearing in both texts in the same order | |
| , | LCS length over candidate length, and over reference length | 0 … 1 |
In words: BLEU is the geometric mean of the n-gram match rates, scaled down for short answers; ROUGE-L is the F1 of the longest shared in-order word sequence.
On a hand example: candidate "a b c d", reference "a c d e": the LCS is "a c d", 3 words, so P = 3/4, R = 3/4, ROUGE-L = 0.75. And "the the the the" against "the cat is here" scores unigram precision 1/4: "the" is credited only as often as the reference contains it.
With the numbers: the Monday answer and the reference are both 11
words (full stop and capitals dropped), so BP = 1. It matches 10 of its 11
words, 8 of its 10 word pairs, 6 of its 9 triples and 4 of its 8 four-word
runs. bleu adds 1 to the top and bottom for n ≥ 2 (smoothing, so one
missing four-word run can't zero the score), giving p = 10/11, 9/11, 7/10,
5/9, and BLEU = exp(¼(ln 0.909 + ln 0.818 + ln 0.7 + ln 0.556)) = 0.73.
Its longest shared in-order run is 10 words, so ROUGE-L = 2 · (10/11) ·
(10/11) / (20/11) = 0.91.
In Python:
import math
ref = "the meeting was moved to friday because the manager is sick".split()
cand = "the meeting was moved to monday because the manager is sick".split()
# every run of n words, in order
def grams(words, n):
return [tuple(words[i:i + n]) for i in range(len(words) - n + 1)]
def p(n):
c, r = grams(cand, n), grams(ref, n)
# clipped: at most as often as ref has it
hits = sum(min(c.count(g), r.count(g)) for g in set(c))
# add-one smoothing for n ≥ 2
s = 1 if n > 1 else 0
return (hits + s) / (len(c) + s)
[round(p(n), 3) for n in range(1, 5)] # → [0.909, 0.818, 0.7, 0.556]
BP = 1.0 if len(cand) > len(ref) else math.exp(1 - len(ref) / len(cand))
# BLEU
round(BP * math.exp(sum(math.log(p(n)) for n in range(1, 5)) / 4), 2) # → 0.73
# every word but "monday", in order
P_LCS, R_LCS = 10 / len(cand), 10 / len(ref)
# ROUGE-L
round(2 * P_LCS * R_LCS / (P_LCS + R_LCS), 2) # → 0.91
# the hand example: LCS "a c d"
P_LCS, R_LCS = 3 / 4, 3 / 4
2 * P_LCS * R_LCS / (P_LCS + R_LCS) # → 0.75
Figure 8 · Drawn from the lesson's code
BLEU and ROUGE-L bars: the correct paraphrase scores lowest while the wrong answer that copies the wording scores almost as high as an exact copy
In code: bleu clips each n-gram count, takes the geometric mean and
applies the brevity penalty; rouge_l finds the longest common subsequence
and returns its F1.
Why it matters in practice. Teams grade open-ended output with an LLM as a judge, a model following an explicit rubric, and embedding-based scores like BERTScore do somewhat better than overlap. But a judge must earn trust first (next section).
Chapter 6
Calibrating a judge: Cohen's kappa
Everyday picture Two teachers grade the same 20 essays pass/fail and agree on 18. Impressive? Not if 18 of the essays were obvious passes: a careful teacher who reads every essay (and passes those 18) and a lazy one who stamps "pass" on everything without reading would also agree on 18. Cohen's kappa subtracts the agreement you'd expect from luck.
Tiny worked example Human labels: 18 pass, 2 fail. A lazy judge says "pass" to everything. Raw agreement is 90%, and chance agreement (both say pass at their own rates) is also 0.9 × 1.0 + 0.1 × 0 = 90%, so kappa = 0. A second example: labels (y, y, n, n) vs. (y, n, n, n) agree 3/4 of the time; chance predicts 0.5 × 0.25 + 0.5 × 0.75 = 0.5; kappa = 0.5.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| kappa: agreement beyond chance | ≤ 1; 0 = chance, 1 = perfect | |
| observed agreement: share of items both raters labelled the same | 0 … 1 | |
| agreement expected by chance | 0 … 1 | |
| one possible label (pass, fail…) | ||
| , | how often rater A (and rater B) uses label ℓ | 0 … 1 |
In words: kappa is how far agreement beats chance, as a share of the most it could beat chance by.
On the worked example: p_o = 0.75, p_e = 0.5, κ = (0.75 − 0.5)/(1 − 0.5) = 0.5.
Level 3: in Python
# rater A's labels
A = ["y", "y", "n", "n"]
# rater B's labels
B = ["y", "n", "n", "n"]
p_o = sum(a == b for a, b in zip(A, B)) / len(A)
# Σ_ℓ p_A(ℓ) p_B(ℓ)
p_e = sum((A.count(ell) / len(A)) * (B.count(ell) / len(B)) for ell in ("y", "n"))
p_o, p_e # → (0.75, 0.5)
# κ
(p_o - p_e) / (1 - p_e) # → 0.5
Figure 9 · Diagram
flowchart LR S[Sample of real outputs] --> H[Humans label<br/>pass / fail] S --> J[LLM judge + rubric<br/>labels pass / fail] H --> K[Cohen's kappa] J --> K K -->|kappa high enough| U[Use the judge at scale] K -->|too low| F[Fix rubric, examples<br/>or judge model] F --> J
In code: cohens_kappa computes p_o and p_e from two lists of labels
and returns κ.
Why it matters in practice. Kappa above about 0.6 is usually considered substantial agreement. Re-check it whenever the judge model or rubric changes.
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: A fraud model has 94% accuracy. Is it good?Think it through, then reveal
A: You can't tell from accuracy. If fraud is 10% of traffic, check recall and precision. In the worked example recall is only 60%, so 4 in 10 frauds slip through.
Question 2Q: When do you optimize for precision vs. recall?Think it through, then reveal
A: By the cost of each error. If a miss is expensive (fraud, cancer screening, a relevant legal document), favor recall. If a false alarm is expensive (blocking a good customer, paging an engineer at 3 a.m.), favor precision. Set the threshold to minimize expected cost.
Question 3Q: What does an AUC of 0.8 mean?Think it through, then reveal
A: A randomly chosen positive gets a higher score than a randomly chosen negative 80% of the time. It measures ranking quality across all thresholds; 0.5 is random.
Question 4Q: Which retrieval metric matters most for RAG, and why?Think it through, then reveal
A: Recall@k, where k is the number of chunks you pass to the model. If the right passage isn't in the top k, no prompt engineering can recover the answer. Add MRR or nDCG when position within the context matters.
Question 5Q: Why not grade LLM answers with BLEU or ROUGE?Think it through, then reveal
A: They measure surface overlap with one reference. A correct paraphrase scores low and a wrong answer that reuses the reference's words scores high. Use code-based checks where possible and a rubric-driven LLM judge otherwise, validated against human labels.
Question 6Q: Your LLM judge agrees with humans 90% of the time. Good enough?Think it through, then reveal
A: Not necessarily. If 90% of answers are "pass", a judge that always says "pass" also agrees 90%. Compute Cohen's kappa to correct for chance agreement.
Primary sources
The papers behind this lesson
Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023): Measured how well strong language models agree with human preferences when used as judges, and catalogued their biases (position, verbosity, self-preference).
Read the annotated companion →The paper ↗Papineni et al., BLEU: a Method for Automatic Evaluation of Machine Translation (2002): Introduced clipped n-gram precision with a brevity penalty.
The paper ↗Lin, ROUGE: A Package for Automatic Evaluation of Summaries (2004): Introduced recall-oriented overlap measures, including the LCS-based ROUGE-L.
The paper ↗Zhang et al., BERTScore: Evaluating Text Generation with BERT (2019): Compared candidate and reference by embedding similarity rather than exact word overlap.
The paper ↗Researcher's shelf
Further reading
- scikit-learn, Metrics and scoring: https://scikit-learn.org/stable/modules/model_evaluation.html
- Receiver operating characteristic (Wikipedia): https://en.wikipedia.org/wiki/Receiver_operating_characteristic
- Discounted cumulative gain (Wikipedia): https://en.wikipedia.org/wiki/Discounted_cumulative_gain
- Cohen's kappa (Wikipedia): https://en.wikipedia.org/wiki/Cohen%27s_kappa
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.