The lesson in one minute
What you'll be able to explain
- Long tasks fail multiplicatively: at 95% per step, 10 steps succeed 60% of the time. Shorten chains and verify at key steps.
- Retry transient failures with capped exponential backoff and jitter; never retry invalid requests; make writes idempotent first.
- Put circuit breakers around dependencies so outages fail fast instead of cascading.
- Detect loops (same call, same arguments) and enforce step and token budgets.
- Most failures are system failures (retrieval, tools, data, evals, adoption), not model failures, and each has a known fix.
Level 1
The practitioner's guide
In one sentence
Most agent failures in production are system failures with known fixes (too many steps, bad retrieval, ambiguous tools, a context that rots, loops, no evals, injected instructions, messy data, brittle integrations, no adoption), and this lesson is the catalogue: symptom, fix, and where the fix is built.
When you need it
When the agent works in demos and fails in production, and nobody can say why. The tell is the gap between the two: demos are three steps long and real tasks are twenty. This lesson's arithmetic explains the gap on its own. If each step succeeds 95% of the time, a ten-step task succeeds 60% of the time and a twenty-step task 36%; at 90% per step, twenty steps succeed 12% of the time; at 99%, 82%. At 95% per step, the most steps you can chain and still succeed nine times in ten is two. Nothing about the model changed between the demo and production; the exponent did. You need this catalogue before the first incident, as a checklist, and after every incident, to name what happened. You don't need the integration patterns (retries, breakers) for a tool that calls nothing outside your process, and you don't need loop detection for a workflow with a fixed number of steps.
Your options
Six defences, from the ones that protect one request to the ones that protect the whole system:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Retries with capped exponential backoff and jitter | Wait 0.1 s, then 0.2 s, doubling to a cap, with a random spread so clients don't retry in waves | A transient blip (a timeout, a rate limit) doesn't fail the task | Latency on the retried request; a duplicate write unless the write is idempotent | Around each outside call |
| Circuit breaker | After a run of failures, calls fail instantly for a cool-down, then one trial call probes for recovery | An outage fails fast and stays contained instead of tying every request up in timeouts (in this lesson's 50-second outage the dead service is called 5 times and 21 requests fail fast) | A fallback path for when the breaker is open | Around each dependency |
| Loop detection and budgets | Stop when the same tool is called with the same arguments three times, or when a step or token budget runs out | A confused agent stops and hands off instead of burning money all night | A false stop now and then; a handoff path | The agent loop |
| Fewer steps, verification, checkpoints | Shorten the chain, check a step's result where it happens and retry only that step, save state to resume from the middle | Attacks the exponent directly: errors stop travelling down the chain | Design effort; a check after each key step costs a call | The task design |
| Contract tests | Automated checks that an outside API still accepts and returns what your tool expects | An upstream change is caught by a test, not a user | Tests to write and keep current | Your test suite |
| Evals and traces | A golden set in CI and a trace of every run | Regressions don't ship silently, and every failure can be replayed to its first bad step | The two lessons before this one | Your pipeline and your monitoring |
How to choose
Read the symptom first; the catalogue maps each one to a fix.
- Long tasks fail and short ones don't: compounding error. Cut steps, verify after the key ones, checkpoint, and put a person at the critical points.
- Timeouts and expired credentials: brittle integrations. Retries with backoff for blips, a breaker for outages, contract tests for changes. Never retry an invalid request (it fails the same way every time), and never retry a write that isn't idempotent.
- The same call repeated, cost spiking: loops. Detect the repeat and set hard budgets.
- Confident wrong answers: bad retrieval, or messy source data. Fix search and ingestion before touching the prompt.
- Wrong tool or bad arguments: ambiguous tools. Fewer tools, precise descriptions, validated inputs.
- The agent obeys text it found in a document: prompt injection. Boundaries around untrusted content and privilege separation.
- It works and nobody uses it: adoption. Build with the people it serves.
- Whatever the symptom, look at the trace for the first failing step before changing anything, and add the case to the golden set.
What it costs
Retries cost time on the slow path: the lesson's waits are 0.1 s and 0.2 s before the third attempt succeeds, capped at 10 s however many attempts there are, and with full jitter each client waits a random time up to that cap. A breaker costs almost nothing while closed and saves a great deal while open: every request that would have waited 30 s on a dead service fails at once instead. Loop detection is a comparison of the last few calls. Verification after a step costs a model call per check, which is why it goes after the key steps, not all of them. The expensive fix is the design one, shortening chains, and it is also the one with the largest payoff: raising per-step reliability from 95% to 99% takes a twenty-step task from 36% to 82%. The lesson's figures draw each of these curves so the trade can be read off, not guessed.
What breaks
- Retrying the invalid. A bad request retried five times is five times the load and the same error. Retry only the exceptions that can clear on their own.
- Retrying a write without an idempotency key. "Create order" twice is two orders.
- Retries in waves. Every client that failed at the same moment retries at the same moment, and the recovering service goes down again. Jitter.
- No breaker. One dead dependency, and every request waits out a full timeout; the whole agent looks down.
- A breaker with no fallback. Failing fast is only useful if the agent can do something else: use a cache, tell the user, hand off.
- Loop detection that only counts. Three calls of the same tool with refining queries is progress, not a loop. Compare arguments, not names.
- Fixing the prompt for a system fault. A retrieval miss, a tool error or a parsing bug looks like a model mistake from the answer alone. The trace shows which it was.
In the wild
The circuit breaker was named by Michael Nygard in
Release It! and written up by Martin Fowler with the three states this
lesson builds (closed, open, half-open); the retry guidance, capped
exponential backoff with jitter, is the AWS Builders' Library article the
lesson links. Libraries package both: in
Python, tenacity gives a retry decorator with exponential and random
exponential waits, stop conditions by attempts or elapsed time, and retry
only on chosen exception types, the same policy as this lesson's
retry_with_backoff. Anthropic's Building effective agents draws the
same conclusions from the other direction: start with simple prompts and
evaluation, add multi-step agents only where simpler designs fall short,
add complexity only when it demonstrably improves outcomes, and have the
agent pause for human feedback at checkpoints; its emphasis on tool design
as an interface problem is the fix for the ambiguous-tools row. The seven
rows this lesson doesn't build point to the lessons that do.
Go deeper
Level 2 builds three of the fixes in plain code: the compounding-error formula and its inverse (how many steps a target allows), retries with capped backoff and full jitter that raise invalid requests at once, a circuit breaker as a three-state machine replayed through an outage, and a loop detector that compares arguments. If you only needed to name the failure and find its fix, you are done.
Level 2
How it works, from scratch
Pilots learn from a catalogue of accidents: each entry names what went wrong, how it showed up in the cockpit, and the checklist item that now prevents it. This lesson is that catalogue for agents: ten failures that keep recurring in production systems, each with its symptom, its fix, and a pointer to the lesson that builds the fix. Three of them get their fix built right here: compounding error (arithmetic), brittle integrations (retries with backoff, and circuit breakers) and runaway loops (loop detection).
| Failure | Symptom | Fix | Lesson |
|---|---|---|---|
| Compounding error | Long tasks fail far more often than short ones | Fewer steps, checks after key steps, checkpoints | primer.agents.planning |
| Bad retrieval | Confident, wrong answers | Hybrid search, reranking, retrieval evals | primer.agents.rag |
| Ambiguous tools | Wrong tool, or bad arguments | Precise descriptions, fewer tools, validation | primer.agents.tools |
| Context rot | Quality drops as a session grows | Summarize, trim, restart with a state handoff | primer.agents.context |
| Loops and runaway | The same call repeated; cost spikes | Step budgets, loop detection, stop conditions | primer.agents.agent_loop |
| No evals | Regressions ship silently | Golden sets in CI, online monitoring | primer.agents.evals |
| Prompt injection | The agent obeys instructions found in data | Untrusted-content boundaries, privilege separation | primer.agents.guardrails |
| Messy enterprise data | Garbled tables, missing permissions | Invest in parsing; permission-aware retrieval | primer.agents.rag |
| Brittle integrations | Timeouts, expired credentials, API changes | Retries with backoff, circuit breakers, contract tests | primer.agents.failures |
| No adoption | It works, and nobody uses it | Build with users, show sources, easy human handoff | primer.agents.deployment |
In code: FAILURES is this table as data, one Failure record per row.
Chapter 1
Compounding error
Everyday picture A relay race where each baton pass succeeds 95% of the time. One pass is nearly safe; ten passes in a row drop the baton more often than you'd think.
Worked example If each step of an agent's task succeeds 95% of the time, independently:
| Steps | Chance every step succeeds |
|---|---|
| 1 | 0.95 |
| 2 | 0.95 × 0.95 = 0.9025 |
| 10 | 0.95¹⁰ ≈ 0.599 |
| 20 | 0.95²⁰ ≈ 0.358 |
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| chance a single step succeeds | 0.95 | |
| number of steps that must all succeed | 10 | |
| multiplied by itself times | 0.599 |
In words: the chance that every step succeeds is the per-step chance multiplied together once per step.
On the worked example: 0.95 multiplied by itself 10 times is 0.599, so a ten-step task at 95% per step fails four times in ten.
Level 3: in Python
p, n = 0.95, 10
# p multiplied by itself n times
round(p ** n, 3) # → 0.599
# fails about four times in ten
round(1 - p ** n, 1) # → 0.4
Figure 2 · Diagram
flowchart LR S1[Step 1<br/>95%] --> S2[Step 2<br/>95%] --> S3[...] --> S10[Step 10<br/>95%] --> D[Done:<br/>60% of the time] S2 -.->|verify, retry<br/>just this step| S2
Figure 1 · Drawn from the lesson's code
Twenty steps succeed 82% of the time at 99% per step but only 36% at 95% and 12% at 90%
In code: chain_success is ; max_steps_for runs it backwards,
returning the most steps you can chain and still reach a target success rate.
Chapter 2
Brittle integrations: retries with backoff
Everyday picture Calling a busy phone line: you don't redial every second. You wait a little, then longer, then longer still, and if everyone redials on the same schedule the line stays jammed, so you add a random pause (jitter).
Worked example A service times out twice and then answers. With a base
delay of 0.1 s the waits are 0.1 s, then 0.2 s, and the third attempt
succeeds. With a cap of 10 s, the waits never exceed 10 s however many
attempts there are. A request that's invalid (a ValueError here) fails
the same way every time, so it's raised immediately: retrying only adds
load.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| wait before retry number | = 0.1 s, = 0.2 s | |
| how many retries have already happened | 0, 1 | |
| base delay | 0.1 s | |
| doubles with every retry | 1, 2, 4, … | |
| the cap on any single wait | 10 s |
In words: each wait doubles the previous one, starting from the base delay, but never exceeds the cap.
On the worked example: = min(10, 0.1 × 1) = 0.1 s and = min(10, 0.1 × 2) = 0.2 s.
Level 3: in Python
b, D_max = 0.1, 10
# d_0 and d_1
[min(D_max, b * 2 ** k) for k in range(2)] # → [0.1, 0.2]
Figure 4 · Diagram
sequenceDiagram participant A as Agent tool participant S as Upstream API A->>S: request S--xA: timeout Note over A: wait 0.1 s A->>S: retry 1 S--xA: timeout Note over A: wait 0.2 s A->>S: retry 2 S-->>A: 200 OK
Figure 3 · Drawn from the lesson's code
The delay doubles each attempt until it hits the 10-second cap, and jitter scatters five clients' retries across the whole range below it
In code: backoff_delays lists the capped waits ;
retry_with_backoff calls a function, retries only the exceptions in
RETRYABLE with those waits (optionally with full jitter), and raises
anything else at once.
Chapter 3
Brittle integrations: circuit breakers
Everyday picture The breaker in your home's fuse box. When a circuit keeps shorting, it trips and cuts the power, instead of letting the wire overheat. After a while you flip it back on to test; if it trips again, it stays off.
A circuit breaker wraps calls to a dependency. After several consecutive failures it opens: calls fail immediately without touching the struggling service (fail fast), so the agent can fall back (use a cache, tell the user, hand off to a person) instead of hanging on timeouts. After a cool-down it goes half-open and lets one trial call through: success closes it, failure opens it again.
Worked example Threshold 3, cool-down 30 s. Three timeouts in a row:
open. A fourth call at 5 s fails instantly with CircuitOpen and the
service isn't called at all. At 31 s, one trial call goes through; it
succeeds, so the breaker closes.
Figure 6 · Diagram
stateDiagram-v2 [*] --> Closed Closed --> Closed: success (reset the failure count) Closed --> Open: 3 failures in a row Open --> Open: calls rejected instantly Open --> HalfOpen: cool-down elapsed HalfOpen --> Closed: trial call succeeds HalfOpen --> Open: trial call fails
Figure 5 · Drawn from the lesson's code
Through a 50-second outage the failing service is called only 5 times, while 21 requests fail fast at the open breaker instead
In code: CircuitBreaker.call is the state diagram: it raises
CircuitOpen while open, lets one trial call through after the cool-down,
and closes on success. simulate_outage replays the outage in the figure.
Contract tests (automated checks that an external API still accepts and returns what your tool expects) catch the third kind of brittleness, upstream API changes, before users do.
Chapter 4
Loops and runaway
Everyday picture A satnav that keeps rerouting you around the same block. The fix isn't a better map; it's noticing that you've passed the same corner three times.
Worked example search("vpn"), search("vpn"), search("vpn"): the
same tool with the same arguments three times in a row. Nothing new can
come back, so the agent is looping. search("vpn"), search("vpn error"),
search("ERR-4012") is the same tool refining its query: progress.
Combine loop detection with hard step and token budgets
(primer.agents.cost.TaskBudget) so a confused agent stops and hands off
instead of burning money.
In code: is_looping reports whether the last few calls were the same
tool with identical arguments.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Your agent works in demos and fails half the time in production. Where do you look first?Think it through, then reveal
At the length of real tasks versus demo tasks (compounding error), and at traces for the first failing step: retrieval misses, wrong tools or arguments, tool errors from integrations, context growth. Then fix per-step reliability where it's lowest, add verification after key steps, and put the failing cases into the eval set.
Question 2Why add jitter to retries?Think it through, then reveal
Without it, every client that failed at the same moment retries at the same moment, producing synchronized waves of load that keep an overloaded service down. Random delays spread the retries out.
Question 3When should you not retry?Think it through, then reveal
When the error can't go away on its own: invalid input, permission denied, not found. And never retry a non-idempotent write without an idempotency key, or you'll duplicate its effect.
Question 4What's the difference between a retry and a circuit breaker?Think it through, then reveal
A retry handles a blip for one request. A breaker handles an outage across many requests: it stops sending traffic to a failing dependency, fails fast so callers can fall back, and probes for recovery.
Primary sources
The papers behind this lesson
Researcher's shelf
Further reading
- Martin Fowler, CircuitBreaker: https://martinfowler.com/bliki/CircuitBreaker.html
- AWS Builders' Library, Timeouts, retries, and backoff with jitter: https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
- Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.