rumblr Work in progressWIP

● The AI Primer · Lesson 53 · Part 2: building systems people rely on

Why the hard ones fail

a catalogue of agent failures, and the fix for each

You'll be able to explain The common failure modes, and the fix for each

Members · open during launch 17 min6 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Long tasks fail multiplicatively: at 95% per step, 10 steps succeed 60% of the time. Shorten chains and verify at key steps.
  2. Retry transient failures with capped exponential backoff and jitter; never retry invalid requests; make writes idempotent first.
  3. Put circuit breakers around dependencies so outages fail fast instead of cascading.
  4. Detect loops (same call, same arguments) and enforce step and token budgets.
  5. Most failures are system failures (retrieval, tools, data, evals, adoption), not model failures, and each has a known fix.

Level 1

The practitioner's guide

In one sentence

Most agent failures in production are system failures with known fixes (too many steps, bad retrieval, ambiguous tools, a context that rots, loops, no evals, injected instructions, messy data, brittle integrations, no adoption), and this lesson is the catalogue: symptom, fix, and where the fix is built.

When you need it

When the agent works in demos and fails in production, and nobody can say why. The tell is the gap between the two: demos are three steps long and real tasks are twenty. This lesson's arithmetic explains the gap on its own. If each step succeeds 95% of the time, a ten-step task succeeds 60% of the time and a twenty-step task 36%; at 90% per step, twenty steps succeed 12% of the time; at 99%, 82%. At 95% per step, the most steps you can chain and still succeed nine times in ten is two. Nothing about the model changed between the demo and production; the exponent did. You need this catalogue before the first incident, as a checklist, and after every incident, to name what happened. You don't need the integration patterns (retries, breakers) for a tool that calls nothing outside your process, and you don't need loop detection for a workflow with a fixed number of steps.

Your options

Six defences, from the ones that protect one request to the ones that protect the whole system:

Option What it does What it guarantees What it costs Where it lives
Retries with capped exponential backoff and jitter Wait 0.1 s, then 0.2 s, doubling to a cap, with a random spread so clients don't retry in waves A transient blip (a timeout, a rate limit) doesn't fail the task Latency on the retried request; a duplicate write unless the write is idempotent Around each outside call
Circuit breaker After a run of failures, calls fail instantly for a cool-down, then one trial call probes for recovery An outage fails fast and stays contained instead of tying every request up in timeouts (in this lesson's 50-second outage the dead service is called 5 times and 21 requests fail fast) A fallback path for when the breaker is open Around each dependency
Loop detection and budgets Stop when the same tool is called with the same arguments three times, or when a step or token budget runs out A confused agent stops and hands off instead of burning money all night A false stop now and then; a handoff path The agent loop
Fewer steps, verification, checkpoints Shorten the chain, check a step's result where it happens and retry only that step, save state to resume from the middle Attacks the exponent directly: errors stop travelling down the chain Design effort; a check after each key step costs a call The task design
Contract tests Automated checks that an outside API still accepts and returns what your tool expects An upstream change is caught by a test, not a user Tests to write and keep current Your test suite
Evals and traces A golden set in CI and a trace of every run Regressions don't ship silently, and every failure can be replayed to its first bad step The two lessons before this one Your pipeline and your monitoring

How to choose

Read the symptom first; the catalogue maps each one to a fix.

  • Long tasks fail and short ones don't: compounding error. Cut steps, verify after the key ones, checkpoint, and put a person at the critical points.
  • Timeouts and expired credentials: brittle integrations. Retries with backoff for blips, a breaker for outages, contract tests for changes. Never retry an invalid request (it fails the same way every time), and never retry a write that isn't idempotent.
  • The same call repeated, cost spiking: loops. Detect the repeat and set hard budgets.
  • Confident wrong answers: bad retrieval, or messy source data. Fix search and ingestion before touching the prompt.
  • Wrong tool or bad arguments: ambiguous tools. Fewer tools, precise descriptions, validated inputs.
  • The agent obeys text it found in a document: prompt injection. Boundaries around untrusted content and privilege separation.
  • It works and nobody uses it: adoption. Build with the people it serves.
  • Whatever the symptom, look at the trace for the first failing step before changing anything, and add the case to the golden set.

What it costs

Retries cost time on the slow path: the lesson's waits are 0.1 s and 0.2 s before the third attempt succeeds, capped at 10 s however many attempts there are, and with full jitter each client waits a random time up to that cap. A breaker costs almost nothing while closed and saves a great deal while open: every request that would have waited 30 s on a dead service fails at once instead. Loop detection is a comparison of the last few calls. Verification after a step costs a model call per check, which is why it goes after the key steps, not all of them. The expensive fix is the design one, shortening chains, and it is also the one with the largest payoff: raising per-step reliability from 95% to 99% takes a twenty-step task from 36% to 82%. The lesson's figures draw each of these curves so the trade can be read off, not guessed.

What breaks

  • Retrying the invalid. A bad request retried five times is five times the load and the same error. Retry only the exceptions that can clear on their own.
  • Retrying a write without an idempotency key. "Create order" twice is two orders.
  • Retries in waves. Every client that failed at the same moment retries at the same moment, and the recovering service goes down again. Jitter.
  • No breaker. One dead dependency, and every request waits out a full timeout; the whole agent looks down.
  • A breaker with no fallback. Failing fast is only useful if the agent can do something else: use a cache, tell the user, hand off.
  • Loop detection that only counts. Three calls of the same tool with refining queries is progress, not a loop. Compare arguments, not names.
  • Fixing the prompt for a system fault. A retrieval miss, a tool error or a parsing bug looks like a model mistake from the answer alone. The trace shows which it was.

In the wild

The circuit breaker was named by Michael Nygard in Release It! and written up by Martin Fowler with the three states this lesson builds (closed, open, half-open); the retry guidance, capped exponential backoff with jitter, is the AWS Builders' Library article the lesson links. Libraries package both: in Python, tenacity gives a retry decorator with exponential and random exponential waits, stop conditions by attempts or elapsed time, and retry only on chosen exception types, the same policy as this lesson's retry_with_backoff. Anthropic's Building effective agents draws the same conclusions from the other direction: start with simple prompts and evaluation, add multi-step agents only where simpler designs fall short, add complexity only when it demonstrably improves outcomes, and have the agent pause for human feedback at checkpoints; its emphasis on tool design as an interface problem is the fix for the ambiguous-tools row. The seven rows this lesson doesn't build point to the lessons that do.

Go deeper

Level 2 builds three of the fixes in plain code: the compounding-error formula and its inverse (how many steps a target allows), retries with capped backoff and full jitter that raise invalid requests at once, a circuit breaker as a three-state machine replayed through an outage, and a loop detector that compares arguments. If you only needed to name the failure and find its fix, you are done.

Level 2

How it works, from scratch

Pilots learn from a catalogue of accidents: each entry names what went wrong, how it showed up in the cockpit, and the checklist item that now prevents it. This lesson is that catalogue for agents: ten failures that keep recurring in production systems, each with its symptom, its fix, and a pointer to the lesson that builds the fix. Three of them get their fix built right here: compounding error (arithmetic), brittle integrations (retries with backoff, and circuit breakers) and runaway loops (loop detection).

Failure Symptom Fix Lesson
Compounding error Long tasks fail far more often than short ones Fewer steps, checks after key steps, checkpoints primer.agents.planning
Bad retrieval Confident, wrong answers Hybrid search, reranking, retrieval evals primer.agents.rag
Ambiguous tools Wrong tool, or bad arguments Precise descriptions, fewer tools, validation primer.agents.tools
Context rot Quality drops as a session grows Summarize, trim, restart with a state handoff primer.agents.context
Loops and runaway The same call repeated; cost spikes Step budgets, loop detection, stop conditions primer.agents.agent_loop
No evals Regressions ship silently Golden sets in CI, online monitoring primer.agents.evals
Prompt injection The agent obeys instructions found in data Untrusted-content boundaries, privilege separation primer.agents.guardrails
Messy enterprise data Garbled tables, missing permissions Invest in parsing; permission-aware retrieval primer.agents.rag
Brittle integrations Timeouts, expired credentials, API changes Retries with backoff, circuit breakers, contract tests primer.agents.failures
No adoption It works, and nobody uses it Build with users, show sources, easy human handoff primer.agents.deployment

In code: FAILURES is this table as data, one Failure record per row.

Chapter 1

Compounding error

Everyday picture A relay race where each baton pass succeeds 95% of the time. One pass is nearly safe; ten passes in a row drop the baton more often than you'd think.

Worked example If each step of an agent's task succeeds 95% of the time, independently:

Steps Chance every step succeeds
1 0.95
2 0.95 × 0.95 = 0.9025
10 0.95¹⁰ ≈ 0.599
20 0.95²⁰ ≈ 0.358
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
chance a single step succeeds 0.95
number of steps that must all succeed 10
multiplied by itself times 0.599

In words: the chance that every step succeeds is the per-step chance multiplied together once per step.

On the worked example: 0.95 multiplied by itself 10 times is 0.599, so a ten-step task at 95% per step fails four times in ten.

Level 3: in Python
p, n = 0.95, 10
# p multiplied by itself n times
round(p ** n, 3)  # → 0.599
# fails about four times in ten
round(1 - p ** n, 1)  # → 0.4

Figure 2 · Diagram

Reading it: every arrow is another chance to drop the baton, so success shrinks with each step. The dotted loop is the remedy: check a step's result where it happens and retry only that step, so an error doesn't travel down the chain. Fewer steps, verification after key steps, checkpoints to resume from the middle, and human review at critical points all attack the same exponent.

Figure 1 · Drawn from the lesson's code

0 5 10 15 20 25 30 steps that must all succeed 0.0 0.2 0.4 0.6 0.8 1.0 chance the whole task succeeds Compounding error 60% 36% 99% per step 95% per step 90% per step

Twenty steps succeed 82% of the time at 99% per step but only 36% at 95% and 12% at 90%

Reading it: the x-axis is the number of steps; each curve is a per-step reliability. At 99% per step a 20-step task still succeeds 82% of the time; at 95% it's 36%; at 90%, 12%. Small gains in per-step reliability are worth far more than they look, and demos with three steps say little about tasks with twenty.

In code: chain_success is ; max_steps_for runs it backwards, returning the most steps you can chain and still reach a target success rate.

Chapter 2

Brittle integrations: retries with backoff

Everyday picture Calling a busy phone line: you don't redial every second. You wait a little, then longer, then longer still, and if everyone redials on the same schedule the line stays jammed, so you add a random pause (jitter).

Worked example A service times out twice and then answers. With a base delay of 0.1 s the waits are 0.1 s, then 0.2 s, and the third attempt succeeds. With a cap of 10 s, the waits never exceed 10 s however many attempts there are. A request that's invalid (a ValueError here) fails the same way every time, so it's raised immediately: retrying only adds load.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
wait before retry number = 0.1 s, = 0.2 s
how many retries have already happened 0, 1
base delay 0.1 s
doubles with every retry 1, 2, 4, …
the cap on any single wait 10 s

In words: each wait doubles the previous one, starting from the base delay, but never exceeds the cap.

On the worked example: = min(10, 0.1 × 1) = 0.1 s and = min(10, 0.1 × 2) = 0.2 s.

Level 3: in Python
b, D_max = 0.1, 10
# d_0 and d_1
[min(D_max, b * 2 ** k) for k in range(2)]  # → [0.1, 0.2]

Figure 4 · Diagram

Reading it: time runs downward. Each failure is followed by a longer pause, which gives an overloaded service room to recover instead of being hammered. Writes must be idempotent (safe to repeat: sending the same request twice has the same effect as once, usually via an idempotency key) before you retry them, or a retried "create order" makes two orders.

Figure 3 · Drawn from the lesson's code

1 2 3 4 5 6 7 retry number 0 2 4 6 8 10 seconds to wait Backoff: base 0.5 s, doubling, capped at 10 s capped exponential (no jitter) five clients with full jitter

The delay doubles each attempt until it hits the 10-second cap, and jitter scatters five clients' retries across the whole range below it

Reading it: the x-axis is the attempt number; the black line is the capped exponential delay. The dots are five clients using full jitter (each waits a random time between zero and the capped delay): instead of all retrying at the same moment, their retries spread out, which is what lets a recovering service actually recover.

In code: backoff_delays lists the capped waits ; retry_with_backoff calls a function, retries only the exceptions in RETRYABLE with those waits (optionally with full jitter), and raises anything else at once.

Chapter 3

Brittle integrations: circuit breakers

Everyday picture The breaker in your home's fuse box. When a circuit keeps shorting, it trips and cuts the power, instead of letting the wire overheat. After a while you flip it back on to test; if it trips again, it stays off.

A circuit breaker wraps calls to a dependency. After several consecutive failures it opens: calls fail immediately without touching the struggling service (fail fast), so the agent can fall back (use a cache, tell the user, hand off to a person) instead of hanging on timeouts. After a cool-down it goes half-open and lets one trial call through: success closes it, failure opens it again.

Worked example Threshold 3, cool-down 30 s. Three timeouts in a row: open. A fourth call at 5 s fails instantly with CircuitOpen and the service isn't called at all. At 31 s, one trial call goes through; it succeeds, so the breaker closes.

Figure 6 · Diagram

Reading it: in the normal state (Closed) calls flow through and failures are counted. Open is the protective state: nothing reaches the dependency. Half-open is a single, cautious test. The breaker turns a slow, cascading failure (every request waiting 30 s on a dead service) into a fast, contained one.

Figure 5 · Drawn from the lesson's code

0 20 40 60 80 100 seconds Circuit breaker (threshold 3, cool-down 15 s) during an outage service down ok failed failed fast (breaker open)

Through a 50-second outage the failing service is called only 5 times, while 21 requests fail fast at the open breaker instead

Reading it: the shaded band is when the upstream service is down; each marker is one request. Before the outage, calls succeed (green). The first three failures (red) open the breaker; after that, requests fail fast (grey) without reaching the service, apart from one trial call per cool-down. Once the service is back, the next trial succeeds and traffic resumes. The service received a handful of calls during the outage instead of all of them.

In code: CircuitBreaker.call is the state diagram: it raises CircuitOpen while open, lets one trial call through after the cool-down, and closes on success. simulate_outage replays the outage in the figure.

Contract tests (automated checks that an external API still accepts and returns what your tool expects) catch the third kind of brittleness, upstream API changes, before users do.

Chapter 4

Loops and runaway

Everyday picture A satnav that keeps rerouting you around the same block. The fix isn't a better map; it's noticing that you've passed the same corner three times.

Worked example search("vpn"), search("vpn"), search("vpn"): the same tool with the same arguments three times in a row. Nothing new can come back, so the agent is looping. search("vpn"), search("vpn error"), search("ERR-4012") is the same tool refining its query: progress. Combine loop detection with hard step and token budgets (primer.agents.cost.TaskBudget) so a confused agent stops and hands off instead of burning money.

In code: is_looping reports whether the last few calls were the same tool with identical arguments.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Your agent works in demos and fails half the time in production. Where do you look first?Think it through, then reveal

At the length of real tasks versus demo tasks (compounding error), and at traces for the first failing step: retrieval misses, wrong tools or arguments, tool errors from integrations, context growth. Then fix per-step reliability where it's lowest, add verification after key steps, and put the failing cases into the eval set.

Question 2Why add jitter to retries?Think it through, then reveal

Without it, every client that failed at the same moment retries at the same moment, producing synchronized waves of load that keep an overloaded service down. Random delays spread the retries out.

Question 3When should you not retry?Think it through, then reveal

When the error can't go away on its own: invalid input, permission denied, not found. And never retry a non-idempotent write without an idempotency key, or you'll duplicate its effect.

Question 4What's the difference between a retry and a circuit breaker?Think it through, then reveal

A retry handles a blip for one request. A breaker handles an outage across many requests: it stops sending traffic to a failing dependency, fails fast so callers can fall back, and probes for recovery.

Primary sources

The papers behind this lesson

Researcher's shelf

Further reading

  • Martin Fowler, CircuitBreaker: https://martinfowler.com/bliki/CircuitBreaker.html
  • AWS Builders' Library, Timeouts, retries, and backoff with jitter: https://aws.amazon.com/builders-library/timeouts-retries-and-backoff-with-jitter/
  • Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.