The lesson in one minute
What you'll be able to explain
- Earn autonomy in steps: shadow mode, then human approval of each action, then autonomy for low-risk actions only; demote on any incident or drift.
- Treat prompts as code: versioned (content hashes), reviewed, evaluated, and one step from rollback.
- Release to a small canary share first and grow it only while its metrics match the current version; roll back automatically otherwise.
- Bound the blast radius: kill switches per tenant and globally, and rate limits on actions.
- Keep a tamper-evident audit log (a hash chain) linking every action to its user, authority and trace.
Level 1
The practitioner's guide
In one sentence
Safe deployment means letting an agent act in the real world one earned step at a time: it proves itself in shadow, then proposes, then acts alone on low-risk work, while every release goes to a few users first, every action can be stopped and rate-limited, and every action is written to a log nobody can quietly edit.
When you need it
The moment an agent's output stops being a draft and
becomes an action: a refund issued, an email sent, a record changed in the
system a company runs on. Offline evals (primer.agents.evals) gate what
ships; this lesson is about limiting the damage from whatever the evals
missed, because real traffic always holds inputs the golden set didn't.
The tell: someone asks "what's the worst this could do in the next five
minutes?" and the answer is "we'd notice eventually". This lesson's worked
example shows why the first step is to watch rather than act: in shadow
mode over ten support tickets the agent agrees with staff 80% of the time
overall, but the split is 100% on replies and 60% on refunds, so the right
first grant of autonomy is replies, not refunds, and only the breakdown
shows it. You don't need canaries and hash chains for an internal tool
that drafts text a person edits; you need all of it for anything that
moves money or changes records, and regulated work needs the audit log
whatever the risk.
Your options
Seven mechanisms, from the ones that cost nothing to run to the ones that hold up under audit. They stack rather than compete:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Shadow mode | The agent records what it would do on real inputs; people keep doing the work | Zero user impact and a clean comparison against what staff actually did, at full volume | Staff carry the whole load meanwhile; you need to log proposals and compare | Your agent loop, plus a comparison job |
| Human approval of each action | The agent proposes; a person approves before anything executes | Nothing executes unseen | Review load, and reviewers who start rubber-stamping | The reviewers' existing workflow |
| Graduated autonomy | Acts alone on low-risk actions once a full window of recent decisions clears a bar; high-risk actions still go to a person; any incident demotes at once | Trust is earned on evidence and taken back on one bad event | Classifying every action's risk; keeping the window's statistics | Your action gate |
| Prompts as code | Every prompt version is named by a hash of its text, reviewed, evaluated, and one step from rollback | Any behaviour change ties to the exact edit that caused it | A registry and the discipline to use it | Your code and configuration |
| Canary release | A new version goes to 1% of users, then 5, 25, 50 and 100, growing only while its success rate stays within a margin of the current version's | At worst a few percent of users saw the bad version for one step | Enough traffic per step to judge (100 canary tasks in this lesson), and metrics split by version | The router and the metrics store |
| Kill switch and rate limit | Stop the agent instantly, per tenant or for everyone; cap actions per minute with a token bucket | A bounded blast radius even when everything else has failed | Someone on call to flip the switch; a burst above the cap is refused | In front of every action |
| Tamper-evident audit log | Each entry stores the hash of the one before, so any edit or deletion breaks every later link | Tampering is detectable; with write-once storage, provable | Storage that can't be modified, and a link from each entry to its trace | Storage |
How to choose
Go by what the action can do if it's wrong, and how fast you'd know.
- Reversible and low value (a reply, a draft, a label): shadow briefly, then approval, then autonomy, with the eval gate and a canary on every prompt or model change.
- Irreversible or high value (a payment, a deletion, an external email): approval stays, whatever the agent's record, above a value threshold you set on purpose. Autonomy applies below it.
- Many customers on one platform: kill switches and rate limits per tenant, so one customer's bad day never becomes everyone's.
- Regulated work: the audit log from day one, on write-once storage, linked to traces, with the agent acting under its own service identity and the narrowest permissions it can do the job with.
- Whatever you pick, promote on a sustained record over a window of decisions (this lesson uses the last 50), never a lucky streak, and demote on a single incident.
What it costs
Shadow mode costs the humans nothing extra and costs you patience: in the lesson's simulation the rolling agreement first clears the 95% bar after 98 decisions, and a window of 50 is the least you can judge on. Approval mode costs reviewer time on every action. A canary costs traffic and time per step; the lesson's controller refuses to judge on fewer than 100 canary tasks, and the SRE Workbook's chapter on canarying makes the trade explicit: a 5% canary with a 20% error rate hurts 1% of requests overall, which is the whole point, but a daily release cycle can't afford a week-long canary. A rate limit costs refused actions during a burst (a bucket of 5 refilling at one per second lets 5 of a burst of 8 through and refuses 3). The audit log costs a hash per entry and storage you can't reclaim. Rollback of a prompt costs nothing, because the previous version is already in the registry. Against all of that, the cost of not having them is the incident: a refund issued on a bad prompt that no switch could stop and no log can show.
What breaks
- Promoting on a streak. Twenty good decisions in a row is luck at 95%. Require a full window.
- A canary judged too early. The lesson's bad version, truly 91% successful, passes the 1% step by luck on 100 tasks and is caught at 5%. Small steps are cheap to get wrong; that's why there are several.
- Users flipping between versions. Assign users by hashing their id into a bucket, so nobody changes version mid-conversation.
- Reviewers who stop reading. Approval mode can bias people toward accepting the agent's suggestion. Sample approvals for a second look, and keep incident demotion automatic.
- Retried writes that double-post. A canary or a retry that re-sends "create order" makes two orders. Use idempotency keys before granting any write.
- An ordinary log table. Anyone with write access can edit a row. Chain the hashes and anchor the latest one somewhere separate.
- A working agent nobody uses. Adoption is part of the job: build with the people whose work it touches, show sources, and make handing a case to a person one click.
In the wild
Canarying comes from site reliability engineering: the Google SRE Workbook defines it as "a partial and time-limited deployment of a change in a service and its evaluation", and its advice to rank metrics by how well they show user-visible problems is why this lesson compares success rates rather than latency alone. Feature flags are the mechanism behind both canaries and kill switches; Martin Fowler's taxonomy separates release, experiment, ops and permissioning toggles, with the kill switch as a long-lived ops toggle, and the pattern is what feature-flag services and OpenFeature, a vendor-neutral specification for feature flagging, exist to implement. The token bucket is a textbook rate-limiting algorithm, linked in Further reading. The hash chain is Haber and Stornetta's 1991 construction for time-stamping documents, the ancestor of every tamper-evident log and of blockchains. The NIST AI Risk Management Framework (released in January 2023) organises this whole lesson's concerns into four functions, Govern, Map, Measure and Manage, and is the vocabulary compliance teams will use for it.
Go deeper
Level 2 builds each gate in plain code: agreement counted
decision by decision, the autonomy state machine with its promotion window,
a content-addressed prompt registry, a canary controller with the rollback
rule worked in numbers, the token bucket formula, and a hash chain you can
break and watch verify() find the break. If you only needed to plan a
rollout, you are done.
Level 2
How it works, from scratch
A new pilot doesn't get the captain's seat on day one. First they sit in the right seat and call out every move they would make while the captain flies. Then they fly with the captain's hands hovering over the controls. Then they fly routine legs alone, while the captain still handles storms. Trust is earned in steps, with evidence at each one, and it can be taken back.
Deploying an agent that takes real actions (refunds, emails, changes to a company's records) works the same way. This lesson builds each mechanism: graduated autonomy, prompt versioning, canary releases, kill switches, rate limits, and a tamper-evident audit log.
Figure 1 · Diagram
flowchart LR
P[Proposed action] --> K{Kill switch<br/>on?}
K -->|yes| X[Blocked]
K -->|no| R{Rate limit<br/>allows it?}
R -->|no| X
R -->|yes| A{Autonomy level<br/>and risk allow<br/>acting alone?}
A -->|no| H[Queue for<br/>human approval]
A -->|yes| E[Execute]
H -->|approved| E
E --> L[Append to<br/>audit log]
In code: each gate is one call: KillSwitch.allowed, then
TokenBucket.allow, then AutonomyController.may_act_alone, and finally
AuditLog.append for whatever executes.
Chapter 1
Graduated autonomy: shadow mode, approval, then autonomy
Everyday picture The trainee pilot again: call out moves (shadow), fly with the captain ready to take over (approval), fly routine legs alone (autonomy for low-risk actions), and hand back the controls after any incident.
In shadow mode the agent runs on real inputs and records what it would do, but takes no action; people keep doing the work. You compare its decisions with theirs. When agreement is high enough, it proposes actions and a person approves each one. When approvals are nearly always "yes", it may act alone on low-risk actions, while high-risk ones keep needing a person. Any incident or drop in quality sends it back a level.
Worked example Ten support tickets in shadow mode:
| Agent proposed | Human did | Agree? |
|---|---|---|
| refund | refund | yes |
| refund | escalate | no |
| reply ×5 | reply ×5 | yes ×5 |
| refund | refund | yes |
| refund | escalate | no |
| refund | refund | yes |
Overall agreement is 8/10 = 80%. Broken down by what the agent proposed, replies agree 5/5 = 100% but refunds only 3/5 = 60%. The breakdown tells you what to let the agent do first: replies, not refunds.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| number of shadow decisions compared | 10 | |
| what the agent proposed for case | refund | |
| what the human actually did for case | escalate | |
| 1 if the condition is true, 0 if not (an indicator) | 0 for case 2 | |
| add up over all cases |
In words: agreement is the share of cases where the agent proposed exactly what the human did.
On the worked example: eight of the ten indicators are 1, so agreement = 8/10 = 0.8.
Level 3: in Python
a = ["refund", "refund"] + ["reply"] * 5 + ["refund", "refund", "refund"]
h = ["refund", "escalate"] + ["reply"] * 5 + ["refund", "escalate", "refund"]
# 𝟙[a_i = h_i]
indicators = [1 if a_i == h_i else 0 for a_i, h_i in zip(a, h)]
indicators # → [1, 0, 1, 1, 1, 1, 1, 1, 0, 1]
# (1/N) Σ over i
sum(indicators) / len(indicators) # → 0.8
Figure 3 · Diagram
stateDiagram-v2 state "Shadow mode: record, don't act" as Shadow state "Human approves each action" as Approve state "Autonomous for low-risk actions" as Auto [*] --> Shadow Shadow --> Approve: agreement ≥ 95% over the last 50 decisions Approve --> Auto: ≥ 98% of the last 50 approved, no incidents Auto --> Approve: incident or quality drift Approve --> Shadow: incident
Figure 2 · Drawn from the lesson's code
Rolling agreement with humans must first clear the 95% bar over a full 50-decision window, and only then is the agent promoted
In code: shadow_agreement computes agreement overall and per proposed
action. AutonomyController is the state diagram: AutonomyController.record_shadow
and AutonomyController.record_approval promote once a full window clears
the bar, while AutonomyController.record_incident and
AutonomyController.record_quality demote at once.
Chapter 2
Prompts are code
Everyday picture A restaurant's recipe binder where every recipe change gets a new revision number, the old version stays in the binder, and the kitchen can switch back in one step if customers complain.
A change to a prompt can break behaviour as badly as a code change, so treat
it the same way: version it, review it, run the evals before release
(primer.agents.evals), and keep the previous version one step away.
PromptRegistry is content-addressed: a version's name includes a
hash of its text. A hash (here SHA-256) is a fixed-length fingerprint:
the same text always gives the same fingerprint, and any change, even one
character, gives a completely different one.
Worked example "Be concise." hashes to ab018bdb…, so its version is
support@ab018bdb. Registering the same text again returns the same
version; "Be concise and friendly." gets a new one. Every trace records
which version produced each model call (primer.agents.observability), so
a behaviour change can be tied to the exact prompt edit that caused it.
In code: PromptRegistry.register hashes the text into a version name
and makes it active; PromptRegistry.rollback steps back one version;
PromptRegistry.active_text returns the prompt currently in use.
Chapter 3
Canary releases: try it on a few users first
Everyday picture Miners once carried a canary into the mine: if the air turned bad, the canary showed it before the miners were harmed. A canary release sends a small share of traffic to the new version, compares its metrics with the current version (the control), and grows the share only while the canary stays healthy.
Worked example Steps 1% → 5% → 25% → 50% → 100%, allowed drop 2 points, at least 100 canary tasks before judging:
| Control success | Canary success | Canary tasks | Decision |
|---|---|---|---|
| 95% | 95% | 100 | advance from 1% to 5% |
| 95% | 90% | 100 | 90 < 95 − 2: roll back to 0% |
| 95% | 90% | 10 | too few samples: stay at 1% |
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| measured success rate on canary traffic (the hat means "measured from samples") | 0.90 | |
| measured success rate on the current version | 0.95 | |
| the largest drop you'll tolerate (Greek delta) | 0.02 |
In words: roll back when the new version's success rate falls more than the allowed margin below the current version's.
On the worked example: 0.90 < 0.95 − 0.02 = 0.93, so roll back.
Level 3: in Python
p_canary, p_control, delta = 0.90, 0.95, 0.02
round(p_control - delta, 2) # → 0.93
# roll back?
p_canary < p_control - delta # → True
Users are assigned to the canary by hashing their id into one of 100 buckets: the same user always lands in the same bucket, so nobody flips between versions mid-conversation, and growing from 5% to 25% only adds users.
Figure 5 · Diagram
sequenceDiagram participant D as Deploy system participant R as Router participant M as Metrics D->>R: send 1% of users to v2 R->>M: success rates, v1 vs v2 M-->>D: v2 95%, v1 95%: healthy D->>R: grow to 5% R->>M: success rates, v1 vs v2 M-->>D: v2 91%, v1 95%: worse by 4 points D->>R: roll back: 0% to v2 D-->>D: alert the owner with traces
Figure 4 · Drawn from the lesson's code
The good version tracks the control through 1, 5, 25, 50 and 100%; the bad one passes the 1% step by luck and is rolled back at 5%
In code: in_canary hashes a user id into one of 100 buckets.
CanaryController.observe accumulates success counts, and
CanaryController.step applies the rollback rule above: advance, wait for
more samples, or roll back. simulate_rollout drives a whole rollout for the
figure.
A feature flag is the same idea as an on/off switch in configuration: it turns a capability on for chosen tenants or users without redeploying, and off again just as fast.
Chapter 4
Kill switches and rate limits: bound the blast radius
Everyday picture The red emergency-stop button on a treadmill, and the turnstile at a stadium that lets people through only so fast however hard the crowd pushes.
A kill switch stops the agent instantly, for one tenant (one customer organisation) or for everyone. A rate limit caps actions per unit time, so even a malfunctioning agent can only send so many emails or change so many records per minute. The blast radius is how much damage a failure can do before someone notices; these two mechanisms bound it.
Worked example: the token bucket. Picture a jar that holds at most 5 tokens and gains 1 token per second. Each action takes a token; with no token, the action is refused. Starting full: 5 actions in a burst succeed, the 6th is refused. After 2 seconds the jar has 2 tokens again: 2 more succeed, the next is refused.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| tokens in the bucket at time | 2 at t = 2 s | |
| when the bucket was last updated | 0 s | |
| tokens left then | 0 after the burst | |
| refill rate, tokens per second | 1 | |
| capacity: the largest burst allowed | 5 | |
| the smaller of the two: the jar can't overflow |
In words: the tokens now are what was left plus what has dripped in since, but never more than the jar holds.
On the worked example: min(5, 0 + 1 × (2 − 0)) = 2 tokens, so two more actions are allowed.
Level 3: in Python
# capacity, tokens per second
C, r = 5, 1
def b(t, t_0, b_t0):
# what was left plus the drip, capped at C
return min(C, b_t0 + r * (t - t_0))
# 2 s after the burst emptied the jar
b(2, 0, 0) # → 2
# a long wait refills only to capacity
b(60, 0, 0) # → 5
Figure 6 · Drawn from the lesson's code
A burst of 8 requests gets 5 through before the bucket empties, then requests pass one per second at the refill rate
In code: KillSwitch holds the global and per-tenant off switches.
TokenBucket.level is the formula for , and TokenBucket.allow
spends a token or refuses.
Chapter 5
A tamper-evident audit log
Everyday picture A ledger where every page ends with a wax seal pressed from the previous page's seal plus this page's contents. Change one number on page 12 and its seal no longer matches, and neither does any seal after it.
An audit log records which agent took which action, for which user, with what inputs. For regulated work, it must be tamper-evident: an alteration can't go unnoticed. A hash chain does this: each entry stores the hash of the previous entry, and its own hash covers its contents and that previous hash.
Level 3: the formula and its symbols
(Entries are numbered from 1 here; verify() reports Python list
positions, which start at 0, so "entry 2" is position 1.)
Symbols
| Symbol | Meaning here |
|---|---|
| the fingerprint (hash) stored with entry | |
| the previous entry's fingerprint | |
| the hash function: any input to a 64-hex-digit fingerprint | |
| "joined together with" (concatenation) | |
| a fixed starting value for the first entry |
In words: each entry's fingerprint is the fingerprint of the previous fingerprint joined to this entry's contents.
On the worked example: three entries: refund A200 for 40, refund A201
for 15, escalate A300. Change entry 2's amount from 15 to 1,500 and
recomputing its fingerprint no longer gives the stored , so verify()
reports position 1. Delete entry 2 instead, and entry 3 moves up to position
1; its stored previous fingerprint is , the deleted entry's, which
doesn't match , so the break is again reported at position 1.
In Python:
import hashlib
def seal(prev, entry):
# SHA256(h_{i-1} ‖ entry_i)
return hashlib.sha256((prev + entry).encode()).hexdigest()
# h_0
log, prev = [], "0" * 64
for entry in ["refund A200 40", "refund A201 15", "escalate A300"]:
prev = seal(prev, entry)
# each entry is stored with its h_i
log.append((entry, prev))
def verify(log):
prev = "0" * 64
for position, (entry, h_i) in enumerate(log):
if seal(prev, entry) != h_i:
# the first broken link
return position
prev = h_i
verify(log) is None # → True
# entry 2's amount changed
verify([log[0], ("refund A201 1500", log[1][1]), log[2]]) # → 1
# entry 2 deleted
verify([log[0], log[2]]) # → 1
Figure 7 · Diagram
flowchart LR G["h0 = 000…"] --> E1["entry 1: refund A200, 40<br/>prev = h0<br/>h1 = SHA256(h0 ‖ entry 1)"] E1 --> E2["entry 2: refund A201, 15<br/>prev = h1<br/>h2 = SHA256(h1 ‖ entry 2)"] E2 --> E3["entry 3: escalate A300<br/>prev = h2<br/>h3 = SHA256(h2 ‖ entry 3)"]
In code: AuditLog.append stores the previous entry's hash in the new
entry and seals it with its own; AuditLog.verify walks the chain and
returns the position of the first broken link, or None.
Chapter 6
Change management: adoption is part of the job
A technically working agent still fails if people don't trust or use it. Involve the people whose work it touches from shadow mode onwards (their corrections are your best eval data), design around their actual workflow, show sources and confidence so they can check answers, and make handing a case to a person one click.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1How would you roll out an agent that takes real actions in a company's ERP system (its finance and operations system of record)?Think it through, then reveal
Start read-only in shadow mode on real transactions, comparing proposals with what staff actually did, broken down by action type. Give the agent a dedicated service identity with the narrowest permissions, a sandbox or dry-run mode for writes, and idempotency keys so retries can't double-post. Move to proposing actions that a person approves in their existing workflow, starting with the low-value, reversible action types that scored best in shadow. Grant autonomy only for those, with value thresholds that still require approval, per-tenant kill switches, rate limits on writes, a canary rollout for every prompt or model change, eval gates in CI, and a tamper-evident audit log tied to traces. Demote automatically on incidents or drift, and keep the finance team involved throughout.
Question 2Why start in shadow mode instead of with human approval?Think it through, then reveal
Shadow mode costs the humans nothing extra and produces a clean comparison against what they actually did, at full volume, before the agent can influence anything. Approval mode adds review load and can bias reviewers toward accepting the agent's suggestion.
Question 3What does a canary protect against that offline evals don't?Think it through, then reveal
Real traffic: inputs, integrations and user behaviour the golden set didn't anticipate. Evals gate the release; the canary limits the damage from whatever the evals missed.
Question 4Why a hash chain rather than an ordinary log table?Think it through, then reveal
Anyone with write access can quietly edit an ordinary row. With a hash chain, any edit or deletion breaks every later link, so tampering is detectable, and anchoring the latest hash somewhere separate (or using write-once storage) makes it provable.
Primary sources
The papers behind this lesson
Haber & Stornetta, How to Time-Stamp a Digital Document (Journal of Cryptology, 1991), Introduced chaining each record to the hash of the one before, the idea behind tamper-evident logs (and, later, blockchains).
The paper ↗Researcher's shelf
Further reading
- Google SRE workbook, canarying releases: https://sre.google/workbook/canarying-releases/
- Martin Fowler, feature toggles: https://martinfowler.com/articles/feature-toggles.html
- Token bucket algorithm: https://en.wikipedia.org/wiki/Token_bucket
- NIST AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.