rumblr Work in progressWIP

● The AI Primer · Lesson 47 · Part 2: building systems people rely on

Multi-step planning

plans, checks, retries and why long tasks fail

You'll be able to explain Plan-and-execute, decomposition, reflection, compounding error

Members · open during launch 19 min6 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Success compounds: . 95% per step is 60% over ten steps.
  2. Plan first, execute step by step, replan from the point of failure, and keep finished work.
  3. Decompose into subtasks with explicit checks; retry only the failed step (checkpoints).
  4. Self-review shares the model's blind spots; tests, schemas and database checks don't.
  5. Retries help only for failures you detect, so the check matters more than the retry.

Level 1

The practitioner's guide

In one sentence

Multi-step planning is how an agent gets a twenty-step job done when each step is only mostly reliable: write the steps down, check each one with something outside the model, save each result so a failure costs one step, and revise the plan from the point of failure rather than starting over.

When you need it

The moment a task takes more than a handful of model calls in a row. The arithmetic is unforgiving: a step that succeeds 95% of the time gives a ten-step task a 60% chance of finishing and a twenty-step task 36%; at fifty steps it is under 8%. Even a 99% step finishes a fifty-step task only 60% of the time. That is why a demo of a three-step task looks great and the same agent on a real twenty-step job disappoints. You don't need any of this for a task that is one or two calls long, or for a fixed sequence your code can run without asking a model what to do next. The tell: an agent that finishes the small cases and fails the long ones, and a bill that shows it restarting from step one after every stumble.

Your options

From the simplest to the most robust; each later row usually keeps the earlier ones:

Option What it does What it guarantees What it costs Where it lives
One step at a time The model decides the next action after each result Adapts to anything it sees Drifts on long tasks; nobody sees the intent before it acts The agent loop
A fixed workflow in code Your code chains the calls in a known order (chaining, routing, parallel branches) Predictable path, easy to test Only for tasks whose steps you know in advance Your code
Plan-and-execute with replanning The model writes the whole plan first, executes it step by step, and rewrites the rest on a failure while keeping finished work A plan a person can read before it runs; recovery that routes around the failure An extra planning call per replan, and a cap on replans Your loop around the model
Decomposition with checks Each subtask has a definition of done that code can check A failure is caught where it happens Writing a check per subtask Your code
Checkpoints Every subtask's output is saved; a retry reruns only the failed step A failure costs one step, not the run Storage for each output Your code, or a workflow engine
External verification with retries Tests, a schema validator, a database query, or a separate grader with a rubric, fed back to the model Concrete failures the model can fix; retries that actually help The check itself, and one more call per retry Outside the model
Human review at critical points A person approves before an expensive or irreversible step Mistakes stop before they cost Latency and attention Your product

How to choose

Start with the simplest thing and add a row only when the numbers say so.

  • Steps you can list in advance: a workflow in code. Anthropic's guidance is to find the simplest solution and add complexity only when needed, reserving agents for open-ended problems where the steps cannot be predicted.
  • Steps the model must discover: plan-and-execute, with replanning and a replan limit so a planner that cannot adapt doesn't loop forever. In this lesson the first plan uses a retired API, the second routes around it, and the step that already succeeded is not run again.
  • Any task longer than a few steps: decompose into subtasks with checks, and checkpoint. The expected work to finish a ten-step task at 95% per step is 10.5 step runs with checkpoints against 13.4 restarting from scratch; at fifty steps it is 53 against about 240.
  • Wherever a wrong step can hide: verify outside the model. Asked to review its own leap-year function, this lesson's model approves the bug; a test against 1900 catches it, and the second draft passes.
  • Wherever a mistake is expensive or irreversible: a person in the loop.
  • Whatever you pick, invest in the check before the retry. With a check that catches every failure and one retry, a 95% step becomes 99.75% and ten steps succeed 97.5% of the time; with a check that catches half, 77%; with no check, 60%. Retries are only as good as the check that triggers them.

What it costs

Planning costs one extra model call up front and one per replan. Checks cost engineering: a test suite, a schema, a query against the source of truth, sometimes a second model with a rubric. Retries cost a call each, but only for failures that were noticed. Checkpoints cost storage and a small amount of bookkeeping, and they are what stop the cost of a long task from growing exponentially: without them every extra step multiplies the expected work by about 1.05, which is a straight line on a log scale. Human review costs latency. Against all of this, the cost of doing nothing is a task that is abandoned and restarted, paying for the same steps over and over.

What breaks

  • A plan reality has invalidated. The API changed, the file moved, the assumption was wrong. Treat the plan as a hypothesis and the first failed step as evidence; replan from there, keeping what is done.
  • Self-review that approves the bug. The reviewer shares the author's blind spots and can confidently reinforce a wrong answer. Put the check outside the model.
  • Retrying a failure nobody detected. A silent failure gets no retry and poisons every later step. Measure your check's detection rate; it matters more than the retry count.
  • Restarting from step one. Without checkpoints a late failure throws away all the work, and long runs become both unreliable and expensive. Save every output; rerun only the failed step.
  • A planner that loops. Cap the replans and the total steps, and report the failure instead of trying forever.
  • Subtasks that cannot be checked. "Reconcile the invoices" has no definition of done; "each invoice appears exactly once in the match" does. Make every subtask small, concrete and verifiable.

In the wild

Wang et al. (2023), Plan-and-Solve Prompting, showed that asking a model to devise a plan and then carry it out cuts the missing-step errors of reasoning straight through. Shinn et al. (2023), Reflexion, turn external feedback such as failed tests into written lessons kept in memory for the next attempt, and report 91% pass@1 on HumanEval against 80% for the base model; the gains come from signals outside the model. Anthropic's Building effective agents names the workflow patterns in the table (prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer) and insists on stopping conditions such as a maximum number of iterations. Temporal is durable execution as a product: it persists a complete event history, and when a process dies another rebuilds the state and resumes where it stopped, with local variables intact. LangGraph's persistence and checkpoints are the same idea inside an agent framework.

Go deeper

Level 2 builds the arithmetic and the remedies in plain Python: the compounding formula and its curves, a plan-and-execute loop that replans from a failed step, five subtasks with checks and the retry that reruns only one of them, the expected-cost formulas with and without checkpoints, and a reflection-versus-verification experiment on a leap-year function. If you only needed to choose, you are done.

Level 2

How it works, from scratch

What follows measures the arithmetic, then builds the three remedies one mechanism at a time.

An agent that does one thing is easy. An agent that does twenty things in a row mostly fails, for a reason that is pure arithmetic. This lesson measures that arithmetic, then builds the three remedies: decompose the task into steps you can check, verify each step with something outside the model, and checkpoint so a failure costs one step instead of the whole run.

Chapter 1

Compounding error: why long tasks fail

Everyday picture A relay race with ten runners. Each runner drops the baton only 1 time in 20 during their leg. That sounds safe, but the team needs all ten legs to go cleanly, and it loses about 4 races in 10.

Tiny worked example Each step of an agent succeeds 95% of the time.

steps chance all succeed
1 0.95
2 0.95 × 0.95 = 0.9025
10 0.95¹⁰ ≈ 0.599
20 0.95²⁰ ≈ 0.358
Level 3: the formula and its symbols

Symbols

Symbol Meaning
chance a single step succeeds
number of steps the task needs

In words: multiply the per-step success rate by itself once per step.

On the example: and . A 95%-reliable step is a 60%-reliable 10-step task.

In Python:

p = 0.95
# ten steps that must all succeed
round(p ** 10, 2)  # → 0.6
round(p ** 20, 2)  # → 0.36

Figure 1 · Drawn from the lesson's code

0 5 10 15 20 25 30 steps in the task (n) 0.0 0.2 0.4 0.6 0.8 1.0 P(task succeeds) = p^n Compounding error: reliable steps, unreliable tasks 90% per step 95% per step 99% per step

At 95% per step a 10-step task succeeds only about 60% of the time, and every curve keeps sliding toward zero as steps are added

Reading it: the x-axis is the number of steps in the task, and the y-axis is the chance the whole task succeeds. Every curve starts near the top and slides down, and the lower the per-step rate, the faster it falls. Find 10 steps on the x-axis and read up to the 95% curve: about 0.6. That's why a demo of a three-step task looks great and the same agent on a real twenty-step job disappoints.

In code: end_to_end_success multiplies the per-step rate by itself once per step, which is the whole formula above.

Why it matters The remedies all attack or : fewer steps (higher-level tools, see primer.agents.tools), checks that catch failures, retries of just the failed step, and human review at critical points.

Chapter 2

Plan-and-execute, with replanning

Everyday picture Cooking from a recipe. You read the whole recipe first (the plan), then do the steps. Halfway through you discover you're out of eggs. You don't start over, and you don't pretend you have eggs. You revise the rest of the recipe with what you've got, keeping everything you've already chopped.

Tiny worked example Goal "Reconcile Q3 invoices". The planner's first plan uses an old API. Here's the actual run:

planner  -> {"steps": ["fetch_payments", "fetch_invoices_v1", "match", "draft_summary"]}
execute  fetch_payments      ok
execute  fetch_invoices_v1   failed: invoices API v1 was retired on 2026-07-01; use fetch_invoices_v2
planner  <- "Completed so far: fetch_payments. Step fetch_invoices_v1 failed: ...retired..."
planner  -> {"steps": ["fetch_invoices_v2", "match", "draft_summary"]}
execute  fetch_invoices_v2   ok
execute  match               ok
execute  draft_summary       ok     (fetch_payments was NOT run again)

Figure 2 · Diagram

Reading it: the inner loop (execute → check → next) is the plan being followed. The outer loop back to Planner only happens on a failed check, and it carries two facts: what's already done (so work is kept) and why the step failed (so the new plan can route around it). The Replans left? diamond stops a planner that can't adapt from looping forever.

The code plan_and_execute(planner, goal) asks for a JSON plan, runs each action, records outputs in state, and on StepFailed asks for a revised plan. Steps already in state are skipped.

In code: PlanRun is what plan_and_execute returns: every plan the planner wrote, each executed step with "ok" or its failure, and the replan count.

Why it matters Deciding one step at a time (a pure ReAct loop, see primer.agents.agent_loop) drifts on long tasks. An upfront plan keeps the agent on track and lets a person see its intent before it acts. The risk is following a plan that reality has invalidated, which is why replanning exists.

Chapter 3

Decomposition into verifiable subtasks

Everyday picture Moving house. "Move house" isn't something you can check off, but "every box labelled", "van booked for Saturday" and "keys handed back" are. Each has a clear definition of done.

Tiny worked example "Reconcile Q3 invoices" becomes five subtasks, each with a check:

subtask output check (definition of done)
fetch_invoices 4 invoices at least one invoice returned
fetch_payments 3 payments at least one payment returned
match 4 rows: invoiced vs paid each invoice appears exactly once
list_mismatches INV-102 (1200 vs 1100), INV-104 (450 vs 0) none
draft_summary "Q3 reconciliation: 4 invoices, 3 payments, 2 mismatches (INV-102 short by 100, INV-104 unpaid)." mentions the mismatches

When the payments API times out once, only fetch_payments runs again: attempts are {fetch_invoices: 1, fetch_payments: 2, match: 1, ...}.

Figure 4 · Diagram

Reading it: arrows show which outputs feed which step. The dotted self-loop on fetch_payments is a retry. Because every step's output is kept (a checkpoint), the retry doesn't redo fetch_invoices.

In code: Subtask pairs a step's work with its check (the definition of done); run_subtasks runs them in order, retries only the step whose check failed, and returns a RunReport of outputs and attempts. reconcile_q3 builds the five subtasks in the table.

Checkpoints: how much work a failure costs

Without checkpoints, a failure anywhere means starting again from step one. With them, you retry just the failed step. The expected number of step executions to finish an -step task:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
chance a single step attempt succeeds (failures are noticed)
number of steps

In words: with checkpoints each step needs on average attempts, so steps need . Restarting from scratch needs successes in a row, and the expected wait for a run of successes grows roughly like .

On the example: , : with checkpoints step runs; restarting from scratch . At it's about 52.6 vs. 240.

In Python:

def with_checkpoints(n, p):
    # each step needs 1/p attempts on average
    return n / p
def restart_from_scratch(n, p):
    return (1 - p ** n) / ((1 - p) * p ** n)
round(with_checkpoints(10, 0.95), 1), round(restart_from_scratch(10, 0.95), 1)  # → (10.5, 13.4)
round(with_checkpoints(50, 0.95), 1), round(restart_from_scratch(50, 0.95))  # → (52.6, 240)

Figure 3 · Drawn from the lesson's code

0 10 20 30 40 50 60 steps in the task (n) 1 0 0 1 0 1 1 0 2 expected step executions (log scale) What a failure costs (p = 0.95 per step) checkpoints: retry the failed step (n/p) no checkpoints: restart from step 1

On the log scale, restarting from scratch climbs as a straight line (exponential growth) while checkpoints stay close to one run per step

Reading it: the x-axis is task length, and the y-axis (log scale) is how many step executions you should expect to pay for. On a log scale each gridline is ten times the one below, so exponential growth draws a straight line. The restart line is that straight line: every extra step multiplies its cost by about . The checkpoint line (, just proportional to ) bends over and flattens, because on a log scale going from 10 to 20 steps rises no more than going from 5 to 10. The widening gap between them is the cost of having no checkpoints, and at 50 steps it's already about 240 step runs against 53. Long tasks without checkpoints aren't just unreliable; they're expensive. Durable execution is the engineering name for this: a workflow engine that saves each step's result so a crashed run resumes where it stopped.

In code: expected_step_runs evaluates both formulas: with checkpoints, the restart formula without.

Chapter 4

Reflection vs. external verification

Everyday picture Proofreading your own essay versus having someone run the numbers in it. You read what you meant to write, so your own blind spots stay blind. A calculator doesn't share them.

Tiny worked example The model writes a leap-year function:

def is_leap(year):
    return year % 4 == 0        # forgets that 1900 was not a leap year

Reflection (asking a model to review it) replies "Looks correct: years divisible by 4 are leap years." External verification (running it against known answers) replies is_leap(1900) returned True, expected False. Fed that failure, the model's second draft passes every case.

Figure 6 · Diagram

Reading it: two paths leave the same first draft. The top path stays inside the model and ends at "approved, still wrong". The bottom path goes through something outside the model (tests, a schema validator, a database query), and that something produces a concrete, checkable failure the model can fix.

In code: self_review is the top path (a model reads the code and approves or not); run_checks is the bottom path (it runs the code against known answers); generate_until_checks_pass loops the bottom path, feeding each failure back to the model until the checks pass.

With verification and retries, the per-step success rate rises:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
chance one attempt succeeds
chance the check detects a failed attempt
retries allowed after a detected failure

In words: you succeed on the first try, or fail and notice and succeed on the next try, and so on. Failures you don't notice get no retry.

On the example: , perfect check , one retry: per step, so ten steps succeed of the time instead of 60%. With a check that catches only half the failures (): .

In Python:

def p_step(p, d, r):
    # p Σ_{i=0}^{r} ((1 - p) d)^i
    return p * sum(((1 - p) * d) ** i for i in range(r + 1))
# a perfect check, one retry
round(p_step(0.95, 1, 1), 4)  # → 0.9975
# ten such steps
round(p_step(0.95, 1, 1) ** 10, 3)  # → 0.975
# a check that catches half the failures
round(p_step(0.95, 0.5, 1), 5)  # → 0.97375
round(p_step(0.95, 0.5, 1) ** 10, 2)  # → 0.77

Figure 5 · Drawn from the lesson's code

0 5 10 15 20 25 30 steps in the task (n) 0.0 0.2 0.4 0.6 0.8 1.0 P(task succeeds) Same 95% step, different verification no checks (step = 0.9500) check catches half, 1 retry (step = 0.9737) check catches all, 1 retry (step = 0.9975) check catches all, 2 retries (step = 0.9999)

With the same 95% step, a ten-step task succeeds 97.5% of the time when a check catches every failure and allows one retry, 77% when the check catches half, and 60% with no checks

Reading it: all curves use the same 95%-reliable step. The bottom curve has no checks. The middle ones add a retry after a failure is caught, with a check that catches half or all failures. The top curve allows two retries. The gap between "half" and "all" is the lesson: retries are only as good as the check that triggers them. Invest in the check.

In code: step_success computes from , and ; feed its result into end_to_end_success to get the ten-step curves.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: Each step of your agent is 95% reliable. How reliable is a 20-step task, and what do you do about it?Think it through, then reveal

A: . Cut the steps (higher-level tools), add external checks after key steps with a retry of just that step, checkpoint so failures don't restart the run, and put human review where a mistake is expensive. With perfect checks and one retry, per-step reliability becomes 99.75%, and 20 steps succeed about 95% of the time.

Question 2Q: Why isn't "ask the model to double-check its work" enough?Think it through, then reveal

A: The reviewer shares the author's blind spots, and self-critique can confidently reinforce a wrong answer. Checks outside the model (tests, schema validation, a query against the source of truth, a separate grader model with a rubric) produce concrete failures the model can act on.

Question 3Q: When is plan-and-execute better than deciding one step at a time?Think it through, then reveal

A: For long, multi-part tasks where staying on track matters and where showing the plan to a person before acting is valuable. Always pair it with replanning. A plan is a hypothesis, and the first failed step is evidence.

Question 4Q: What makes a subtask "good"?Think it through, then reveal

A: It's small, concrete and verifiable. It has a definition of done that code can check, and its output is saved so a later failure doesn't redo it.

Primary sources

The papers behind this lesson

Wang et al., Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models (2023).

It showed that asking a model to first devise a plan and then carry it out step by step reduces missed steps compared with reasoning straight through.

The paper ↗
Shinn et al., Reflexion: Language Agents with Verbal Reinforcement Learning (2023).

Agents that turn feedback from the environment (failed tests, wrong answers) into written lessons and retry improve markedly, and the gains come from external signals.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
  • Lilian Weng, LLM Powered Autonomous Agents (planning, reflection): https://lilianweng.github.io/posts/2023-06-23-agent/
  • Temporal, durable execution: https://docs.temporal.io/
  • LangGraph docs (persistence and checkpoints): https://langchain-ai.github.io/langgraph/

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.