rumblr Work in progressWIP

● The AI Primer · Lesson 52 · Part 2: building systems people rely on

Safe deployment

letting an agent act in the real world, one earned step at a time

You'll be able to explain Shadow mode, graduated autonomy, canaries, kill switches, audit logs

Members · open during launch 23 min7 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Earn autonomy in steps: shadow mode, then human approval of each action, then autonomy for low-risk actions only; demote on any incident or drift.
  2. Treat prompts as code: versioned (content hashes), reviewed, evaluated, and one step from rollback.
  3. Release to a small canary share first and grow it only while its metrics match the current version; roll back automatically otherwise.
  4. Bound the blast radius: kill switches per tenant and globally, and rate limits on actions.
  5. Keep a tamper-evident audit log (a hash chain) linking every action to its user, authority and trace.

Level 1

The practitioner's guide

In one sentence

Safe deployment means letting an agent act in the real world one earned step at a time: it proves itself in shadow, then proposes, then acts alone on low-risk work, while every release goes to a few users first, every action can be stopped and rate-limited, and every action is written to a log nobody can quietly edit.

When you need it

The moment an agent's output stops being a draft and becomes an action: a refund issued, an email sent, a record changed in the system a company runs on. Offline evals (primer.agents.evals) gate what ships; this lesson is about limiting the damage from whatever the evals missed, because real traffic always holds inputs the golden set didn't. The tell: someone asks "what's the worst this could do in the next five minutes?" and the answer is "we'd notice eventually". This lesson's worked example shows why the first step is to watch rather than act: in shadow mode over ten support tickets the agent agrees with staff 80% of the time overall, but the split is 100% on replies and 60% on refunds, so the right first grant of autonomy is replies, not refunds, and only the breakdown shows it. You don't need canaries and hash chains for an internal tool that drafts text a person edits; you need all of it for anything that moves money or changes records, and regulated work needs the audit log whatever the risk.

Your options

Seven mechanisms, from the ones that cost nothing to run to the ones that hold up under audit. They stack rather than compete:

Option What it does What it guarantees What it costs Where it lives
Shadow mode The agent records what it would do on real inputs; people keep doing the work Zero user impact and a clean comparison against what staff actually did, at full volume Staff carry the whole load meanwhile; you need to log proposals and compare Your agent loop, plus a comparison job
Human approval of each action The agent proposes; a person approves before anything executes Nothing executes unseen Review load, and reviewers who start rubber-stamping The reviewers' existing workflow
Graduated autonomy Acts alone on low-risk actions once a full window of recent decisions clears a bar; high-risk actions still go to a person; any incident demotes at once Trust is earned on evidence and taken back on one bad event Classifying every action's risk; keeping the window's statistics Your action gate
Prompts as code Every prompt version is named by a hash of its text, reviewed, evaluated, and one step from rollback Any behaviour change ties to the exact edit that caused it A registry and the discipline to use it Your code and configuration
Canary release A new version goes to 1% of users, then 5, 25, 50 and 100, growing only while its success rate stays within a margin of the current version's At worst a few percent of users saw the bad version for one step Enough traffic per step to judge (100 canary tasks in this lesson), and metrics split by version The router and the metrics store
Kill switch and rate limit Stop the agent instantly, per tenant or for everyone; cap actions per minute with a token bucket A bounded blast radius even when everything else has failed Someone on call to flip the switch; a burst above the cap is refused In front of every action
Tamper-evident audit log Each entry stores the hash of the one before, so any edit or deletion breaks every later link Tampering is detectable; with write-once storage, provable Storage that can't be modified, and a link from each entry to its trace Storage

How to choose

Go by what the action can do if it's wrong, and how fast you'd know.

  • Reversible and low value (a reply, a draft, a label): shadow briefly, then approval, then autonomy, with the eval gate and a canary on every prompt or model change.
  • Irreversible or high value (a payment, a deletion, an external email): approval stays, whatever the agent's record, above a value threshold you set on purpose. Autonomy applies below it.
  • Many customers on one platform: kill switches and rate limits per tenant, so one customer's bad day never becomes everyone's.
  • Regulated work: the audit log from day one, on write-once storage, linked to traces, with the agent acting under its own service identity and the narrowest permissions it can do the job with.
  • Whatever you pick, promote on a sustained record over a window of decisions (this lesson uses the last 50), never a lucky streak, and demote on a single incident.

What it costs

Shadow mode costs the humans nothing extra and costs you patience: in the lesson's simulation the rolling agreement first clears the 95% bar after 98 decisions, and a window of 50 is the least you can judge on. Approval mode costs reviewer time on every action. A canary costs traffic and time per step; the lesson's controller refuses to judge on fewer than 100 canary tasks, and the SRE Workbook's chapter on canarying makes the trade explicit: a 5% canary with a 20% error rate hurts 1% of requests overall, which is the whole point, but a daily release cycle can't afford a week-long canary. A rate limit costs refused actions during a burst (a bucket of 5 refilling at one per second lets 5 of a burst of 8 through and refuses 3). The audit log costs a hash per entry and storage you can't reclaim. Rollback of a prompt costs nothing, because the previous version is already in the registry. Against all of that, the cost of not having them is the incident: a refund issued on a bad prompt that no switch could stop and no log can show.

What breaks

  • Promoting on a streak. Twenty good decisions in a row is luck at 95%. Require a full window.
  • A canary judged too early. The lesson's bad version, truly 91% successful, passes the 1% step by luck on 100 tasks and is caught at 5%. Small steps are cheap to get wrong; that's why there are several.
  • Users flipping between versions. Assign users by hashing their id into a bucket, so nobody changes version mid-conversation.
  • Reviewers who stop reading. Approval mode can bias people toward accepting the agent's suggestion. Sample approvals for a second look, and keep incident demotion automatic.
  • Retried writes that double-post. A canary or a retry that re-sends "create order" makes two orders. Use idempotency keys before granting any write.
  • An ordinary log table. Anyone with write access can edit a row. Chain the hashes and anchor the latest one somewhere separate.
  • A working agent nobody uses. Adoption is part of the job: build with the people whose work it touches, show sources, and make handing a case to a person one click.

In the wild

Canarying comes from site reliability engineering: the Google SRE Workbook defines it as "a partial and time-limited deployment of a change in a service and its evaluation", and its advice to rank metrics by how well they show user-visible problems is why this lesson compares success rates rather than latency alone. Feature flags are the mechanism behind both canaries and kill switches; Martin Fowler's taxonomy separates release, experiment, ops and permissioning toggles, with the kill switch as a long-lived ops toggle, and the pattern is what feature-flag services and OpenFeature, a vendor-neutral specification for feature flagging, exist to implement. The token bucket is a textbook rate-limiting algorithm, linked in Further reading. The hash chain is Haber and Stornetta's 1991 construction for time-stamping documents, the ancestor of every tamper-evident log and of blockchains. The NIST AI Risk Management Framework (released in January 2023) organises this whole lesson's concerns into four functions, Govern, Map, Measure and Manage, and is the vocabulary compliance teams will use for it.

Go deeper

Level 2 builds each gate in plain code: agreement counted decision by decision, the autonomy state machine with its promotion window, a content-addressed prompt registry, a canary controller with the rollback rule worked in numbers, the token bucket formula, and a hash chain you can break and watch verify() find the break. If you only needed to plan a rollout, you are done.

Level 2

How it works, from scratch

A new pilot doesn't get the captain's seat on day one. First they sit in the right seat and call out every move they would make while the captain flies. Then they fly with the captain's hands hovering over the controls. Then they fly routine legs alone, while the captain still handles storms. Trust is earned in steps, with evidence at each one, and it can be taken back.

Deploying an agent that takes real actions (refunds, emails, changes to a company's records) works the same way. This lesson builds each mechanism: graduated autonomy, prompt versioning, canary releases, kill switches, rate limits, and a tamper-evident audit log.

Figure 1 · Diagram

Reading it: every action the agent proposes passes the same gates in the same order, cheapest and most absolute first. The kill switch stops everything instantly; the rate limit bounds how much damage even a misbehaving agent can do per minute; the autonomy level decides whether a person must approve. Whatever executes is written to the audit log.

In code: each gate is one call: KillSwitch.allowed, then TokenBucket.allow, then AutonomyController.may_act_alone, and finally AuditLog.append for whatever executes.

Chapter 1

Graduated autonomy: shadow mode, approval, then autonomy

Everyday picture The trainee pilot again: call out moves (shadow), fly with the captain ready to take over (approval), fly routine legs alone (autonomy for low-risk actions), and hand back the controls after any incident.

In shadow mode the agent runs on real inputs and records what it would do, but takes no action; people keep doing the work. You compare its decisions with theirs. When agreement is high enough, it proposes actions and a person approves each one. When approvals are nearly always "yes", it may act alone on low-risk actions, while high-risk ones keep needing a person. Any incident or drop in quality sends it back a level.

Worked example Ten support tickets in shadow mode:

Agent proposed Human did Agree?
refund refund yes
refund escalate no
reply ×5 reply ×5 yes ×5
refund refund yes
refund escalate no
refund refund yes

Overall agreement is 8/10 = 80%. Broken down by what the agent proposed, replies agree 5/5 = 100% but refunds only 3/5 = 60%. The breakdown tells you what to let the agent do first: replies, not refunds.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
number of shadow decisions compared 10
what the agent proposed for case refund
what the human actually did for case escalate
1 if the condition is true, 0 if not (an indicator) 0 for case 2
add up over all cases

In words: agreement is the share of cases where the agent proposed exactly what the human did.

On the worked example: eight of the ten indicators are 1, so agreement = 8/10 = 0.8.

Level 3: in Python
a = ["refund", "refund"] + ["reply"] * 5 + ["refund", "refund", "refund"]
h = ["refund", "escalate"] + ["reply"] * 5 + ["refund", "escalate", "refund"]
# 𝟙[a_i = h_i]
indicators = [1 if a_i == h_i else 0 for a_i, h_i in zip(a, h)]
indicators  # → [1, 0, 1, 1, 1, 1, 1, 1, 0, 1]
# (1/N) Σ over i
sum(indicators) / len(indicators)  # → 0.8

Figure 3 · Diagram

Reading it: each box is a level of trust; each arrow is labelled with the evidence needed to move. Promotions need a sustained record over a window of recent decisions, not a lucky streak; demotions happen at once, on a single incident. Even at the top level, high-risk actions (large refunds, deletions) keep needing a person.

Figure 2 · Drawn from the lesson's code

0 25 50 75 100 125 150 175 200 shadow decisions so far 0.6 0.7 0.8 0.9 1.0 agreement with humans Shadow mode: earning the first promotion window not yet full agreement, last 50 decisions promotion bar (95%) promoted to approval mode

Rolling agreement with humans must first clear the 95% bar over a full 50-decision window, and only then is the agent promoted

Reading it: the x-axis counts shadow decisions; the blue line is agreement with humans over the last 50 of them. The dashed line is the 95% bar. Early on the window is too short to judge (grey region); the marker shows where the rolling agreement first clears the bar and the agent is promoted to proposing actions for approval.

In code: shadow_agreement computes agreement overall and per proposed action. AutonomyController is the state diagram: AutonomyController.record_shadow and AutonomyController.record_approval promote once a full window clears the bar, while AutonomyController.record_incident and AutonomyController.record_quality demote at once.

Chapter 2

Prompts are code

Everyday picture A restaurant's recipe binder where every recipe change gets a new revision number, the old version stays in the binder, and the kitchen can switch back in one step if customers complain.

A change to a prompt can break behaviour as badly as a code change, so treat it the same way: version it, review it, run the evals before release (primer.agents.evals), and keep the previous version one step away. PromptRegistry is content-addressed: a version's name includes a hash of its text. A hash (here SHA-256) is a fixed-length fingerprint: the same text always gives the same fingerprint, and any change, even one character, gives a completely different one.

Worked example "Be concise." hashes to ab018bdb…, so its version is support@ab018bdb. Registering the same text again returns the same version; "Be concise and friendly." gets a new one. Every trace records which version produced each model call (primer.agents.observability), so a behaviour change can be tied to the exact prompt edit that caused it.

In code: PromptRegistry.register hashes the text into a version name and makes it active; PromptRegistry.rollback steps back one version; PromptRegistry.active_text returns the prompt currently in use.

Chapter 3

Canary releases: try it on a few users first

Everyday picture Miners once carried a canary into the mine: if the air turned bad, the canary showed it before the miners were harmed. A canary release sends a small share of traffic to the new version, compares its metrics with the current version (the control), and grows the share only while the canary stays healthy.

Worked example Steps 1% → 5% → 25% → 50% → 100%, allowed drop 2 points, at least 100 canary tasks before judging:

Control success Canary success Canary tasks Decision
95% 95% 100 advance from 1% to 5%
95% 90% 100 90 < 95 − 2: roll back to 0%
95% 90% 10 too few samples: stay at 1%
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
measured success rate on canary traffic (the hat means "measured from samples") 0.90
measured success rate on the current version 0.95
the largest drop you'll tolerate (Greek delta) 0.02

In words: roll back when the new version's success rate falls more than the allowed margin below the current version's.

On the worked example: 0.90 < 0.95 − 0.02 = 0.93, so roll back.

Level 3: in Python
p_canary, p_control, delta = 0.90, 0.95, 0.02
round(p_control - delta, 2)  # → 0.93
# roll back?
p_canary < p_control - delta  # → True

Users are assigned to the canary by hashing their id into one of 100 buckets: the same user always lands in the same bucket, so nobody flips between versions mid-conversation, and growing from 5% to 25% only adds users.

Figure 5 · Diagram

Reading it: the deploy system only ever acts on measurements. Each growth step waits for enough traffic to judge, and a single bad reading sends everyone back to the old version automatically. At worst, a few percent of users saw the bad version for one step.

Figure 4 · Drawn from the lesson's code

1% 5% 25% 50% 100% share of traffic on the canary 0.92 0.93 0.94 0.95 0.96 0.97 task success rate good version (95%) 1% 5% share of traffic on the canary bad version (91%) control canary tolerance (2 pts) rolled back to 0%

The good version tracks the control through 1, 5, 25, 50 and 100%; the bad one passes the 1% step by luck and is rolled back at 5%

Reading it: each panel is one rollout. The x-axis is the rollout step (with the share of traffic on the canary); the lines are measured success rates for control and canary. The good version (left) tracks the control and reaches 100%. The bad version (right), truly 91% successful, gets lucky on the 100 tasks of the 1% step and advances, then falls below the tolerance band at 5% and is rolled back to 0%, marked by the red cross. Small steps are cheap to get wrong; that's why there are several.

In code: in_canary hashes a user id into one of 100 buckets. CanaryController.observe accumulates success counts, and CanaryController.step applies the rollback rule above: advance, wait for more samples, or roll back. simulate_rollout drives a whole rollout for the figure.

A feature flag is the same idea as an on/off switch in configuration: it turns a capability on for chosen tenants or users without redeploying, and off again just as fast.

Chapter 4

Kill switches and rate limits: bound the blast radius

Everyday picture The red emergency-stop button on a treadmill, and the turnstile at a stadium that lets people through only so fast however hard the crowd pushes.

A kill switch stops the agent instantly, for one tenant (one customer organisation) or for everyone. A rate limit caps actions per unit time, so even a malfunctioning agent can only send so many emails or change so many records per minute. The blast radius is how much damage a failure can do before someone notices; these two mechanisms bound it.

Worked example: the token bucket. Picture a jar that holds at most 5 tokens and gains 1 token per second. Each action takes a token; with no token, the action is refused. Starting full: 5 actions in a burst succeed, the 6th is refused. After 2 seconds the jar has 2 tokens again: 2 more succeed, the next is refused.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
tokens in the bucket at time 2 at t = 2 s
when the bucket was last updated 0 s
tokens left then 0 after the burst
refill rate, tokens per second 1
capacity: the largest burst allowed 5
the smaller of the two: the jar can't overflow

In words: the tokens now are what was left plus what has dripped in since, but never more than the jar holds.

On the worked example: min(5, 0 + 1 × (2 − 0)) = 2 tokens, so two more actions are allowed.

Level 3: in Python
# capacity, tokens per second
C, r = 5, 1
def b(t, t_0, b_t0):
    # what was left plus the drip, capped at C
    return min(C, b_t0 + r * (t - t_0))
# 2 s after the burst emptied the jar
b(2, 0, 0)  # → 2
# a long wait refills only to capacity
b(60, 0, 0)  # → 5

Figure 6 · Drawn from the lesson's code

0 1 2 3 4 5 6 seconds 0 1 2 3 4 5 tokens Token bucket: capacity 5, refill 1 per second tokens in the bucket allowed refused

A burst of 8 requests gets 5 through before the bucket empties, then requests pass one per second at the refill rate

Reading it: the blue line is the number of tokens in the jar over time; green dots are allowed actions and red crosses refused ones. The opening burst drains the jar, refusals follow, and then actions trickle through at the refill rate: bursts are allowed, sustained floods are not.

In code: KillSwitch holds the global and per-tenant off switches. TokenBucket.level is the formula for , and TokenBucket.allow spends a token or refuses.

Chapter 5

A tamper-evident audit log

Everyday picture A ledger where every page ends with a wax seal pressed from the previous page's seal plus this page's contents. Change one number on page 12 and its seal no longer matches, and neither does any seal after it.

An audit log records which agent took which action, for which user, with what inputs. For regulated work, it must be tamper-evident: an alteration can't go unnoticed. A hash chain does this: each entry stores the hash of the previous entry, and its own hash covers its contents and that previous hash.

Level 3: the formula and its symbols

(Entries are numbered from 1 here; verify() reports Python list positions, which start at 0, so "entry 2" is position 1.)

Symbols

Symbol Meaning here
the fingerprint (hash) stored with entry
the previous entry's fingerprint
the hash function: any input to a 64-hex-digit fingerprint
"joined together with" (concatenation)
a fixed starting value for the first entry

In words: each entry's fingerprint is the fingerprint of the previous fingerprint joined to this entry's contents.

On the worked example: three entries: refund A200 for 40, refund A201 for 15, escalate A300. Change entry 2's amount from 15 to 1,500 and recomputing its fingerprint no longer gives the stored , so verify() reports position 1. Delete entry 2 instead, and entry 3 moves up to position 1; its stored previous fingerprint is , the deleted entry's, which doesn't match , so the break is again reported at position 1.

In Python:

import hashlib
def seal(prev, entry):
    # SHA256(h_{i-1} ‖ entry_i)
    return hashlib.sha256((prev + entry).encode()).hexdigest()
# h_0
log, prev = [], "0" * 64
for entry in ["refund A200 40", "refund A201 15", "escalate A300"]:
    prev = seal(prev, entry)
    # each entry is stored with its h_i
    log.append((entry, prev))
def verify(log):
    prev = "0" * 64
    for position, (entry, h_i) in enumerate(log):
        if seal(prev, entry) != h_i:
            # the first broken link
            return position
        prev = h_i
verify(log) is None  # → True
# entry 2's amount changed
verify([log[0], ("refund A201 1500", log[1][1]), log[2]])  # → 1
# entry 2 deleted
verify([log[0], log[2]])  # → 1

Figure 7 · Diagram

Reading it: each box carries the fingerprint of the box before it, so the boxes are chained. Editing any box changes its fingerprint, which breaks the link to the next box, so verification walks the chain and stops at the first broken link. In production, also write the log to storage that can't be modified afterwards (write-once storage) and link each entry to its full trace.

In code: AuditLog.append stores the previous entry's hash in the new entry and seals it with its own; AuditLog.verify walks the chain and returns the position of the first broken link, or None.

Chapter 6

Change management: adoption is part of the job

A technically working agent still fails if people don't trust or use it. Involve the people whose work it touches from shadow mode onwards (their corrections are your best eval data), design around their actual workflow, show sources and confidence so they can check answers, and make handing a case to a person one click.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1How would you roll out an agent that takes real actions in a company's ERP system (its finance and operations system of record)?Think it through, then reveal

Start read-only in shadow mode on real transactions, comparing proposals with what staff actually did, broken down by action type. Give the agent a dedicated service identity with the narrowest permissions, a sandbox or dry-run mode for writes, and idempotency keys so retries can't double-post. Move to proposing actions that a person approves in their existing workflow, starting with the low-value, reversible action types that scored best in shadow. Grant autonomy only for those, with value thresholds that still require approval, per-tenant kill switches, rate limits on writes, a canary rollout for every prompt or model change, eval gates in CI, and a tamper-evident audit log tied to traces. Demote automatically on incidents or drift, and keep the finance team involved throughout.

Question 2Why start in shadow mode instead of with human approval?Think it through, then reveal

Shadow mode costs the humans nothing extra and produces a clean comparison against what they actually did, at full volume, before the agent can influence anything. Approval mode adds review load and can bias reviewers toward accepting the agent's suggestion.

Question 3What does a canary protect against that offline evals don't?Think it through, then reveal

Real traffic: inputs, integrations and user behaviour the golden set didn't anticipate. Evals gate the release; the canary limits the damage from whatever the evals missed.

Question 4Why a hash chain rather than an ordinary log table?Think it through, then reveal

Anyone with write access can quietly edit an ordinary row. With a hash chain, any edit or deletion breaks every later link, so tampering is detectable, and anchoring the latest hash somewhere separate (or using write-once storage) makes it provable.

Primary sources

The papers behind this lesson

Haber & Stornetta, How to Time-Stamp a Digital Document (Journal of Cryptology, 1991), Introduced chaining each record to the hash of the one before, the idea behind tamper-evident logs (and, later, blockchains).

The paper ↗

Researcher's shelf

Further reading

  • Google SRE workbook, canarying releases: https://sre.google/workbook/canarying-releases/
  • Martin Fowler, feature toggles: https://martinfowler.com/articles/feature-toggles.html
  • Token bucket algorithm: https://en.wikipedia.org/wiki/Token_bucket
  • NIST AI Risk Management Framework: https://www.nist.gov/itl/ai-risk-management-framework

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.