At a glance
Key takeaways
- An eval is a fixed set of real tasks plus graders; rerun it on every prompt, model, tool or retrieval change.
- Grade outcomes and end state, plus constraints on the path (required and forbidden tools, step limits), not an exact sequence of calls.
- Use code graders wherever possible; use an LLM judge with a rubric for open-ended output, and calibrate it against humans with Cohen's kappa.
- Report cost per successful task and p95 latency beside quality.
- Gate releases on regressions, and turn every production failure into a golden task.
Level 2
How it works, from scratch
A driving test doesn't ask the learner to describe driving; it puts them on a fixed route with known hazards and a checklist. Pass the route, get the licence. Change the car, retake the test.
An eval (evaluation) is that test for an AI system: a fixed set of real tasks, a way to judge each result, and a score. Every time you change a prompt, a model, a tool or a retrieval setting, you rerun it. For agents this is the core engineering discipline: without it, every change is a guess, and regressions reach users silently.
Figure 1 · Diagram
flowchart LR
G[Golden set<br/>real tasks + expected outcomes] --> R[Run the agent<br/>version under test]
R --> T[Trajectories:<br/>answer, tool calls,<br/>end state, tokens, time]
T --> C[Code graders<br/>exact, schema, end state]
T --> J[LLM judge<br/>rubric, calibrated]
C --> S[Scorecard<br/>success, steps, cost, p95]
J --> S
S --> D{Better than the<br/>current version?}
D -->|yes| SHIP[Ship]
D -->|no| FIX[Fix and rerun]
In code: evaluate is one lap of this diagram: it runs every golden task
through run_task, grades each run and returns the scorecard.
Chapter 1
The golden set: a fixed route with known answers
Everyday picture A teacher's answer key, built from past exam papers that real students actually got wrong.
Worked example This module's golden set has four support tasks against a small order database (A100 shipped, total 25; A200 delivered, total 40; A300 delivered, total 900; refunds above 500 must be escalated to a person):
| Task id | User says | Pass means |
|---|---|---|
status |
"Where is order A100?" | answer names A100 and "shipped"; no refund issued |
refund-small |
"Please refund order A200, it arrived broken." | A200 ends up refunded |
refund-over-limit |
"Refund order A300, the TV was damaged." | A300 ends up escalated, not refunded |
out-of-scope |
"Can you change my shipping address to Paris?" | answer hands off to the support team; no refund |
Start with even 50 cases taken from real traffic, and grow the set by adding every production failure you find (see "Closing the loop").
In code: Task holds one golden case (the prompt, the expected end state,
required and forbidden tools, a step limit), and GOLDEN_TASKS is the table
above. Run holds everything one attempt left behind: answer, tool calls,
end state, tokens, time and cost.
Chapter 2
Grade outcomes, not paths
Everyday picture A maths teacher who marks the answer and checks that you didn't use a calculator, but doesn't insist you solved it with their favourite method.
An agent can reach the right result by many routes, so you grade the end state and constraints, not the exact sequence of steps:
- Outcome: is the database in the expected end state? Does the answer contain the facts it must?
- Constraints on the path (the trajectory): were required tools used, were forbidden tools avoided, did it stay under the step limit?
Worked example For refund-small, a run that calls refund first and
lookup_order second passes: the end state is right, the required tool was
used, and 2 steps ≤ 5. For status, a run with a perfect answer that also
called refund fails: refund is forbidden on a status question.
Figure 2 · Diagram
flowchart TD
R[A finished run] --> E{End state as<br/>expected?}
E -->|no| F[Fail]
E -->|yes| A{Answer has the<br/>required facts?}
A -->|no| F
A -->|yes| Q{Required tools used,<br/>forbidden tools avoided?}
Q -->|no| F
Q -->|yes| S{Steps within<br/>the limit?}
S -->|no| F
S -->|yes| P[Pass]
Code graders (exact match, schema validation, database checks) are cheap, fast and deterministic: use them wherever the outcome can be checked by code. Trajectory metrics measure how well it got there:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here |
|---|---|
| the set of tools this task needs | |
| the set of tools the agent actually called | |
| tools in both sets | |
| how many tools a set contains |
In words: the share of the tools the task needs that the agent actually used.
On a worked example: a refund needs {lookup_order, refund}; an agent that only looked the order up scores |{lookup_order}| / 2 = 0.5.
In Python:
T_required = {"lookup_order", "refund"}
T_called = {"lookup_order"}
# ∩: tools in both sets
T_required & T_called # → {'lookup_order'}
len(T_required & T_called) / len(T_required) # → 0.5
In code: grade_run walks the diamonds in the diagram and returns a
Grade with every reason a run failed; exact_match is the simplest code
grader; tool_selection_accuracy is the formula above.
Chapter 3
LLM-as-judge, and checking the judge
Everyday picture A teaching assistant grades a big stack of essays using the professor's rubric. Before trusting the TA's grades, the professor grades 20 of the same essays and compares.
For open-ended answers, code can't decide "good". An LLM judge is a model given an explicit rubric (a short list of pass/fail criteria) and asked to grade. It's only useful if it agrees with people, so you calibrate it: have humans label a sample and measure agreement.
Raw agreement is misleading, because a judge that always says "pass" agrees with humans on every answer humans passed. Cohen's kappa corrects for the agreement you'd get by chance:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Range |
|---|---|---|
| kappa (Greek letter), the chance-corrected agreement | 1 = perfect, 0 = no better than chance, below 0 = worse | |
| observed agreement: share of items where judge and human give the same label | 0 … 1 | |
| agreement expected by chance, if each labelled at their own rates independently | 0 … 1 | |
| a label, here "pass" or "fail" | ||
| share of items the human gave label | 0 … 1 | |
| add up over both labels |
In words: kappa is how much of the possible improvement over chance agreement the judge actually achieves.
On the worked example: 20 answers; both say pass on 10, both fail on 6, they disagree on 4. So = 16/20 = 0.80. The human passes 12/20 = 0.6 and the judge passes 12/20 = 0.6, so = 0.6·0.6 + 0.4·0.4 = 0.52. Then κ = (0.80 − 0.52) / (1 − 0.52) = 0.28 / 0.48 = 0.583. A judge that always says pass agrees 60% of the time, but = 0.6·1 + 0.4·0 = 0.6, so κ = (0.6 − 0.6) / 0.4 = 0: all of its agreement is luck. A common rule of thumb is to want κ ≥ 0.6 before trusting a judge unsupervised; this one, at 0.583, is flagged for a better rubric.
Level 3: in Python
# both pass on 10, both fail on 6
p_o = (10 + 6) / 20
p_human = {"pass": 12 / 20, "fail": 8 / 20}
p_judge = {"pass": 12 / 20, "fail": 8 / 20}
# Σ over the labels
p_e = sum(p_human[label] * p_judge[label] for label in p_human)
round(p_e, 2) # → 0.52
# κ
round((p_o - p_e) / (1 - p_e), 3) # → 0.583
always_pass = {"pass": 1.0, "fail": 0.0}
p_e = sum(p_human[label] * always_pass[label] for label in p_human)
# agrees 60% of the time, all by luck
round((0.6 - p_e) / (1 - p_e), 3) # → 0.0
Figure 3 · Diagram
sequenceDiagram
participant H as Human reviewers
participant J as LLM judge
participant C as Calibration
H->>C: labels for 20 sampled answers
J->>C: labels for the same 20
C->>C: agreement p_o, chance p_e, kappa
alt kappa >= 0.6
C-->>J: trusted for this rubric and judge model
else kappa < 0.6
C-->>H: fix the rubric, add examples, re-check
end
Figure 4 · Drawn from the lesson's code
Near-perfect judge: 95% agreement, kappa 0.89; worked example: 80%, kappa 0.58; coin flip at 50% and always-pass at 60% both score kappa 0
In code: rubric_judge asks a judge model to grade an answer against the
rubric; cohens_kappa computes from two lists of labels;
calibrate_judge reports agreement and kappa and decides whether the judge is
trusted. always_pass_judge is the useless judge from the example.
Chapter 4
Metrics for retrieval-augmented answers
For RAG (retrieval-augmented generation; see primer.agents.rag), measure
the two halves separately. Retrieval recall@k: did the right document
appear in the top k results? Faithfulness: what share of the answer's
claims are supported by the retrieved sources (computed here with
primer.agents.guardrails.groundedness)? If recall is low, fix search; if
recall is high but faithfulness is low, fix generation.
In code: recall_at_k scores the retrieval half; faithfulness scores
the generation half as the share of supported claims.
Chapter 5
Quality next to cost and speed
Always report cost and latency beside quality: a change that adds 2% success but doubles cost may not be worth it. Latency is reported as p95, the 95th percentile: sort all response times, and p95 is the time that 95% of requests beat. Averages hide the slow tail that users notice. The key cost number is cost per successful task (total cost ÷ number of successes), because a cheap run that fails still has to be paid for.
In code: percentile computes p95 by the sort-and-count rule above;
evaluate puts p95 latency and cost per successful task on the scorecard
beside the success rate.
Chapter 6
The release gate: catching regressions
Everyday picture Changing the recipe of a best-selling cake: before selling the new version, you bake both and check that nothing the customers liked got worse.
A regression is a task that used to pass and now fails. The gate compares the candidate with the current version on the whole golden set and blocks the release if any task regressed or the success rate dropped.
Worked example Version v1 passes all four tasks. Version v2's prompt was
edited to "always be maximally helpful", and it now refunds A300 (900, over
the 500 limit) instead of escalating. v2 is also cheaper, because it skips
the lookup. The gate reports regressed_tasks = ["refund-over-limit"] and
blocks the release: cheaper and wrong.
Figure 5 · Diagram
sequenceDiagram participant Dev as Engineer participant CI as CI pipeline participant E as Eval runner Dev->>CI: change the prompt (v2) CI->>E: run golden set on v1 and v2 E-->>CI: v1: 4/4 pass, v2: 3/4 pass CI->>CI: task refund-over-limit passed before, fails now CI-->>Dev: release blocked, with the failing trace
Figure 6 · Drawn from the lesson's code
v1 passes all four golden tasks; v2 passes three and fails only the over-limit refund
Figure 7 · Drawn from the lesson's code
v2 makes one tool call on each refund where v1 makes two, so it is slightly cheaper per success (2.24 vs 2.38 dollars per 1,000) despite the broken limit
In code: run_task runs either version against a fresh copy of the order
database; compare_versions is the gate: it lists the regressed tasks and
blocks the release on any regression or success-rate drop.
Chapter 7
Online evaluation: closing the loop
Offline sets miss what real users do. In production, collect explicit
feedback (thumbs up or down) and implicit signals: rephrasing the same
question, retrying, abandoning, or asking for a human all suggest the answer
didn't help. Sample live traces (the recorded step-by-step history of a request; see
primer.agents.observability) for human review, alert on drift (a
metric creeping away from its usual level after a deploy or as traffic
changes), and turn
every confirmed failure into a new golden task.
Figure 8 · Diagram
flowchart LR P[Production traffic] --> S[Signals<br/>thumbs, rephrase,<br/>retry, abandon, escalate] S --> T[Flagged traces] T --> H[Human review] H --> G[New golden task] G --> CI[Release gate] CI --> P
In code: a Signal is one piece of user feedback;
implicit_dissatisfaction_rate is the share of sessions with any unhappy
signal; promote_to_golden turns a confirmed failure into a new Task.
Test yourself
5 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1How do you evaluate an agent whose correct path isn't fixed?Think it through, then reveal
Grade what must be true at the end, not how it got there: the end state of the systems it touched (the database row, the ticket, the file), the facts the answer must contain, and constraints on the path (required tools used, forbidden tools avoided, step and cost limits). Add trajectory metrics (tool selection accuracy, argument correctness, steps, tokens) as diagnostics. For open-ended outputs, use a rubric-based LLM judge that's calibrated against human labels.
Question 2Your LLM judge agrees with humans 85% of the time. Is it good?Think it through, then reveal
Not necessarily. If 85% of answers are good, a judge that always says "pass" also agrees 85% of the time. Compute Cohen's kappa, which subtracts chance agreement; look at the disagreements; and recheck whenever the judge model or rubric changes.
Question 3A new prompt improves the average score. Do you ship it?Think it through, then reveal
Only after checking for regressions task by task, cost and latency changes, and whether the improvement is bigger than run-to-run noise (rerun, or use more cases). An average can go up while a critical case, like an over-limit refund, breaks.
Question 4How do you build the first eval set for a new agent?Think it through, then reveal
Collect 30 to 50 real requests (from logs, support tickets, or domain experts), write down for each what must be true afterwards, and write code checks for as many as possible. Run it on every change from day one. Then grow it from production: every flagged failure becomes a case.
Question 5What would you monitor once the agent is live?Think it through, then reveal
Task success (from outcome checks and sampled human review), implicit dissatisfaction (rephrase, retry, abandon, escalate rates), cost per successful task, p95 latency, tool error rates, and drift in any of these after a deploy.
Primary sources
The papers behind this lesson
Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023), Measured how well strong models agree with human preferences when used as graders, and catalogued their biases (position, verbosity, self-preference), the basis for calibrating a judge before trusting it.
Read the annotated companion →The paper ↗Cohen, A Coefficient of Agreement for Nominal Scales (1960), Introduced kappa, agreement corrected for chance, which this lesson uses to check a judge.
The paper ↗Es, James, Espinosa-Anke & Schockaert, RAGAS: Automated Evaluation of Retrieval Augmented Generation (2023), Proposed reference-free metrics such as faithfulness and context relevance that score the retrieval and generation halves of a RAG system separately.
The paper ↗Researcher's shelf
Further reading
- Anthropic, Define success criteria and build evaluations: https://docs.claude.com/en/docs/test-and-evaluate/develop-tests
- Anthropic, Demystifying evals for AI agents: https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023): https://arxiv.org/abs/2306.05685
- RAGAS documentation (RAG metrics): https://docs.ragas.io/
- Cohen's kappa: https://en.wikipedia.org/wiki/Cohen%27s_kappa
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.