The lesson in one minute
What you'll be able to explain
- Autonomy spectrum: fixed workflow → router → agent loop → multi-agent. Use the least that works.
- Named patterns: chaining (with gates), routing (with a fallback), parallelization (sectioning, voting), orchestrator-workers, evaluator-optimizer.
- Multi-agent buys focused context and parallelism, and costs tokens, hand-off losses and debuggability.
- Production shape: an explicit state machine, the model consulted only at specific nodes, checkpoints after each step.
Level 1
The practitioner's guide
In one sentence
Orchestration is the decision of how much of a system's control flow the model gets to decide, from none (code runs fixed steps and asks the model to fill each one) to all of it (a team of models steering each other), and the named patterns in between.
When you need it
You face this decision the moment a task needs more than one model call: a document pipeline, a support desk, a research assistant, anything that must look something up and then act. The tell that you have chosen wrongly in one direction is a fixed pipeline that keeps growing special cases for requests it didn't anticipate; the tell in the other is an agent that spends a minute and a thousand tokens on a request a two-step script answered correctly every time. Anthropic's Building effective agents, the post these pattern names come from, draws the line this way: workflows are model calls "orchestrated through predefined code paths"; agents are systems where the model "dynamically directs its own processes and tool usage". It also gives the rule: find the simplest solution possible, and add complexity only when needed. You don't need orchestration at all when one call with retrieval and a few examples in the prompt answers the question; that post says many applications stop there.
This lesson measures the cost of each step up on one question ("How many PTO days do I get, and what is the meal per-diem?"). A fixed workflow answers it in 2 calls and 93 tokens; a router in 2 calls and 105 tokens; an agent loop in 2 calls and about 990 tokens, ten times the workflow, because it re-sends its tool definitions and results on every call; a supervisor with two specialists in 4 calls and about 600 tokens. Same answer, four prices.
Your options
Seven designs, from the least autonomy to the most:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Fixed workflow (prompt chaining with gates) | Code runs steps in order; the model fills each one; a code check between steps stops a bad result | The same steps every time; a failed gate stops before the next call is paid for or misled | 2 calls and 93 tokens on the lesson's question; a gate per hand-off to write | Your code |
| Router | One cheap classifier call picks a handler; unknown labels go to a fallback | Exactly one path per request, and never a crash on a label the model invented | One extra call (105 tokens here) and a fallback handler | Your code, with one model decision |
| Parallelization (sectioning, voting) | Independent pieces run at once, or several judges answer the same prompt and the majority wins | Wall-clock time of the slowest branch; three 80%-accurate judges make an 89.6% panel if they err independently | n times the tokens for n branches | Your code (a thread pool) |
| Evaluator-optimizer | One call drafts, another critiques, until it passes or the round limit hits | A checked result, or a flagged best effort at the limit | Up to two calls per round (2 rounds in the lesson's example) | Your code |
| Agent loop | The model picks tools and decides when it is done | Handles requests you did not anticipate | About 990 tokens for the lesson's question, ten times the workflow; less predictability | The model decides; your loop enforces the budget (primer.agents.agent_loop) |
| Orchestrator-workers, multi-agent | A lead model splits the work, specialists each get only their brief, a synthesizer combines | Focused context per specialist and parallel work; subtasks decided at run time | 4 calls and about 600 tokens here; in Anthropic's production research system, about 15 times the tokens of a chat | Several model roles, coordinated by your code |
| Explicit state machine with checkpoints | Code owns every state and transition; the model is consulted at a few named nodes; the run is saved after each step | Auditable, testable transitions and a crash that costs one step, not the run | A state model to design and a store for checkpoints | Your code, or a durable-execution engine |
How to choose
Start from whether the steps vary between requests, or only what happens inside each step.
- Steps known in advance, variation inside them (invoice intake: classify, extract, validate, approve, post): a fixed workflow or state machine, with the model at the judgement nodes only.
- A few kinds of request, each with its own handling: a router in front of the workflows, with a fallback that is safe rather than clever.
- One hard judgement call: voting, but vary the prompt, model or evidence, because copies of the same model tend to make the same mistake.
- Quality that a checker can state as criteria (a citation present, a test passing): evaluator-optimizer, with a round limit.
- The path depends on what is discovered along the way: an agent loop.
- Work that is separable and too big for one context (independent research threads, per-document processing): orchestrator-workers. Not for tasks where every part needs the same context or depends on the others; the same Anthropic report found most coding tasks fall on that side.
- Whatever you pick, code owns the transitions you need to audit, and the model decides only where judgement is needed.
What it costs
Tokens, first: each rung up the spectrum multiplies them, from 93 to 105 to about 990 to about 600 on one toy question, and Anthropic reports agents using about 4 times the tokens of a chat and multi-agent systems about 15 times. Latency: a chain is the sum of its calls, a parallel fan-out is its slowest branch, and every hand-off in a multi-agent design adds a call. Quality: multi-agent can win (that report measured a 90.2% improvement over a single agent on an internal research evaluation) when the work parallelizes, and loses context at every hand-off when it does not. Effort: every pattern in this lesson is under 100 lines of plain Python; a framework adds tracing, persistence and approval hooks, and hidden assumptions. Crashes: without checkpoints, a rerun after a crash repeats every model call (4 instead of 2 in the lesson's invoice run) and risks a different answer the second time.
What breaks
- A router that trusts the label. The classifier returns free text
(
"Billing-ish, maybe refunds?"). Normalize it, accept only known labels, send the rest to a fallback, and log the misses. - A gate that isn't there. Without a check between chain steps, a one-line outline becomes a paid-for, confident draft of nothing.
- Correlated voters. Five copies of one model with one prompt agree on the same error; the majority formula assumes independence.
- An evaluator that is never satisfied. No round limit, no exit. Put the limit in, and flag what came out at the limit.
- Lost hand-offs. A specialist knows only what the supervisor wrote in its brief. Whatever the supervisor forgot is gone. Test that each specialist sees only its own task, and that a delegation to a specialist that doesn't exist is reported, not crashed on.
- No checkpoint. A crash on the last step restarts from the first, doubling cost and, combined with a non-idempotent side effect, posting an invoice twice.
In the wild
The pattern names come from Anthropic's Building effective agents, which also advises starting with the model API directly before reaching for a framework. LangGraph models a workflow as a graph with explicit state and checkpoints, which is the state-machine row above. The Claude Agent SDK and the OpenAI Agents SDK ship vendor-built agent loops, tools and hand-offs. CrewAI and AutoGen are built around multi-agent teams. Temporal provides durable execution for any code: each completed step's result is stored and a crashed run resumes from there. Anthropic's own research feature is an orchestrator-workers system with a lead model and parallel subagents, and its engineering report is the source of the token multipliers above.
Go deeper
Level 2 builds every pattern in a few dozen lines each and measures it: the four-rung cost table, a chain with a gate, a router with a fallback, voting with the formula for why independent judges help, a supervisor whose specialists see only their brief, and an invoice state machine that survives a crash. If you only needed to choose, you are done.
Level 2
How it works, from scratch
"Agent" covers everything from one model call inside ordinary code to a team of models steering each other. This lesson lays out that range, builds each named pattern in a few dozen lines, and measures what each step up costs. The rule that falls out is simple: use the least autonomy that solves the problem, and say so out loud.
Chapter 1
The autonomy spectrum
Everyday picture Four ways to run an office. An assembly line (fixed workflow): the steps never change and people do one task at each station. A receptionist (router) decides which department you go to. A personal assistant (agent loop) decides which errands to run and when the job's done. A team with a manager (multi-agent) splits the work among specialists.
Figure 2 · Diagram
flowchart LR W[Fixed workflow<br/>code decides steps] --> R[Router<br/>LLM picks a path] R --> A[Agent loop<br/>LLM picks tools] A --> M[Multi-agent<br/>agents coordinate]
Tiny worked example One question, "How many PTO days do I get, and
what is the meal per-diem when travelling?", answered by each design
(autonomy_costs()):
| design | model calls | why |
|---|---|---|
| fixed workflow | 2 | rewrite as a search, then answer |
| router | 2 | classify, then the chosen handler answers |
| agent loop | 2 | one turn with two parallel tool calls, one answer turn |
| multi-agent | 4 | plan, one call per specialist (2), final answer |
Figure 1 · Drawn from the lesson's code
Workflow, router and agent loop each make 2 calls and multi-agent 4, but the agent loop uses about 990 tokens, ten times the workflow's 93
Chapter 2
The named patterns
These names (from Anthropic's Building effective agents) let you describe a design in one word.
In code: every pattern below is built from ask, one model call with
one user message that returns the reply's text.
Prompt chaining
Everyday picture A relay race with an inspector at each hand-off: if the baton is dropped, the next runner doesn't start.
Tiny worked example Step 1 writes an outline, and a code gate checks it has at
least three points. The outline "- scope" fails the gate, so the chain stops after one
call and the draft step never runs.
Figure 4 · Diagram
flowchart LR
I[Input] --> S1[LLM: outline] --> G{Gate: 3+ points?}
G -->|yes| S2[LLM: draft] --> O[Output]
G -->|no| X[Stop with the reason]
In code: run_chain runs a list of ChainSteps in order, feeding each
output into the next prompt and stopping at the first gate that reports a
problem; ChainResult records the outputs and where and why it stopped.
Routing
Everyday picture The receptionist listens for five seconds and sends you to billing, tech support or the front desk.
Tiny worked example The classifier replies " Technical\n", which is
normalized to technical and routed. It replies "Billing-ish, maybe refunds?",
which is not a known label, so the request goes to the general fallback
instead of crashing.
Figure 5 · Diagram
flowchart LR Q[Request] --> C[LLM classifier] C -->|billing| B[Billing handler] C -->|technical| T[Tech handler] C -->|anything else| F[Fallback]
In code: route makes the one classifier call, normalizes the label,
swaps anything unknown for the fallback, and hands the request to that
label's handler.
Parallelization: sectioning and voting
Everyday picture Sectioning: several cooks each make one dish at the same time. Voting: a panel of judges, where the majority decides.
Tiny worked example Three reviewers judge a SQL snippet: two say
"vulnerable" and one says "safe", so the verdict is ("vulnerable", {"vulnerable": 2, "safe": 1}).
Figure 6 · Diagram
flowchart LR
Q[Same prompt] --> R1[Reviewer 1]
Q --> R2[Reviewer 2]
Q --> R3[Reviewer 3]
R1 --> V{Majority}
R2 --> V
R3 --> V
Why voting helps, when voters err independently:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning |
|---|---|
| number of voters (odd, so there are no ties) | |
| chance a single voter is right | |
| number of voters who are right | |
| ways to choose which of the are right |
In words: add up the chances of every outcome where more than half the voters are right.
On the example: , : . Three 80% judges make an 89.6% panel.
In Python:
from math import comb
n, p = 3, 0.8
# the chance exactly k voters are right ...
P = sum(comb(n, k) * p ** k * (1 - p) ** (n - k)
# ... for every k from ⌊n/2⌋ + 1 to n
for k in range(n // 2 + 1, n + 1))
round(P, 3) # → 0.896
Figure 3 · Drawn from the lesson's code
Majority accuracy rises with more independent voters: at 70% per voter, 5 voters reach 84% and 15 reach 95%; at 60% per voter the climb is slow
In code: run_sections is sectioning: it runs independent pieces of
work on a thread pool and collects every result. vote is voting: it asks
each reviewer the same prompt through run_sections and returns the
majority label with its tally, and majority_accuracy evaluates the formula
above.
Orchestrator-workers
Everyday picture A project lead reads the brief, splits it into sections, hands each to a writer, then edits the pieces into one report.
Tiny worked example "What's the meal per-diem, when are receipts due,
and does PTO roll over?" The orchestrator returns three subtasks. Three workers each
find one document (fin-002, fin-004, hr-001), and the synthesizer
writes one answer citing all three.
Figure 7 · Diagram
flowchart TD Q[Request] --> O[Orchestrator LLM:<br/>decide the subtasks] O --> W1[Worker: per-diem] O --> W2[Worker: receipts] O --> W3[Worker: PTO rollover] W1 --> S[Synthesizer LLM:<br/>one cited answer] W2 --> S W3 --> S
In code: orchestrate asks the orchestrator for subtasks, runs the
workers on them in parallel with run_sections, and asks the synthesizer for
one answer, returning an OrchestratorResult. policy_worker is the worker
in the example: it finds the one current policy document for a subtask.
Evaluator-optimizer
Everyday picture A writer and an editor. The draft goes back and forth until the editor signs off, or the deadline hits.
Tiny worked example Draft 1, "PTO rolls over.", gets the feedback "Missing a citation". Draft 2, "Up to 5 unused PTO days roll over [hr-001].", passes. Result: 2 rounds, passed.
Figure 8 · Diagram
flowchart LR
T[Task] --> G[Generator drafts]
G --> E{Evaluator:<br/>PASS?}
E -->|feedback| G
E -->|PASS| D[Done]
E -->|max rounds| B[Stop: best effort, flagged]
In code: evaluate_optimize alternates generator drafts and evaluator
verdicts, feeding each critique back, and returns the last draft, the rounds
used and whether it passed.
Chapter 3
Multi-agent: a supervisor and specialists
Everyday picture A manager who never does the work, only assigns it and assembles the results. Each specialist gets a short, focused brief.
Tiny worked example The supervisor sends "PTO days per year?" to hr and
"Meal per-diem?" to finance. The HR specialist never sees the finance question,
and vice versa (tested). A delegation to a non-existent legal specialist
is reported, not crashed on.
Figure 9 · Diagram
sequenceDiagram participant U as User participant S as Supervisor participant H as HR agent participant F as Finance agent U->>S: PTO and meal allowance? S->>H: "PTO days per year?" (only this) S->>F: "Meal per-diem?" (only this) H-->>S: 20 days [hr-001] F-->>S: 75 dollars a day [fin-002] S-->>U: combined answer
When it's worth it. Use several agents when subtasks are genuinely separable, such as independent research threads or per-document work that would overflow one context. The costs are real: more tokens, context lost at every hand-off, and much harder debugging.
In code: supervise gets a delegation plan from the supervisor, sends
each specialist only its own task (reporting unknown names instead of
crashing), then asks the supervisor for the final answer; SupervisorResult
keeps the delegations and every specialist's answer.
Chapter 4
An explicit state machine, with the model at specific nodes
Everyday picture A board game. The squares and the rules for moving between them are fixed. On a few special squares you draw a card (ask the model). Everywhere else, the rules decide.
Tiny worked example Invoice intake. The model is asked twice: "is this an invoice?" and "extract the number, vendor and amount". Code does everything else: validation, the approval rule (over 5,000 waits for a person), and posting. The runs:
INVOICE INV-2041 ... 1200.00 -> RECEIVED > CLASSIFIED > EXTRACTED > VALIDATED > POSTED > DONE
INVOICE INV-2042 ... 18000.00 -> RECEIVED > CLASSIFIED > EXTRACTED > VALIDATED > AWAITING_APPROVAL
Acme newsletter -> RECEIVED > REJECTED
Figure 11 · Diagram
stateDiagram-v2 [*] --> RECEIVED RECEIVED --> CLASSIFIED: LLM says invoice RECEIVED --> REJECTED: LLM says other CLASSIFIED --> EXTRACTED: LLM extracts fields EXTRACTED --> VALIDATED: code checks pass EXTRACTED --> NEEDS_REVIEW: code checks fail VALIDATED --> AWAITING_APPROVAL: amount > limit VALIDATED --> POSTED: code posts to ledger AWAITING_APPROVAL --> POSTED: human approves POSTED --> DONE
In code: InvoiceWorkflow.run walks the states, calling the model only
at RECEIVED and CLASSIFIED; validate_invoice is the code check between
EXTRACTED and VALIDATED; InvoiceWorkflow.approve is the human step that
releases a paused invoice. WorkflowRun holds the current state, its history
and the extracted data.
Checkpoints and durable execution. After every transition the run is saved to disk. Durable execution means a workflow whose progress survives crashes, because each completed step's result is stored and the run resumes from there. Engines like Temporal do this at scale.
Figure 10 · Drawn from the lesson's code
After one crash while posting, resuming from the checkpoint finishes in 2 model calls, while restarting from scratch takes 4
In code: InvoiceWorkflow saves a checkpoint file after every
transition and loads it at the start of InvoiceWorkflow.run, so a rerun
resumes. resume_comparison crashes the ledger once and counts the model
calls each way.
Chapter 5
Frameworks vs. plain code
LangGraph models agents as graphs with explicit state and checkpoints. The Claude Agent SDK and OpenAI Agents SDK give you vendor-built agent loops, tools and hand-offs. CrewAI and AutoGen focus on multi-agent teams. Temporal provides durable execution for any code. A framework earns its place with standard patterns, built-in tracing, persistence and human-in-the-loop hooks. Plain code wins for simple flows, full control and fewer dependencies. Everything in this lesson is under 100 lines of plain Python, and knowing that is what lets you judge a framework.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: When would you use a fixed workflow instead of an agent? Give an example.Think it through, then reveal
A: When the steps are known in advance and the variation is inside each step, not in which steps happen. Invoice intake is the example: classify, extract, validate, approve, post. A workflow is cheaper, faster, testable and auditable. The model is used where judgement is needed (classification, extraction), and code owns the flow. Reach for an agent only when the path genuinely depends on what's discovered along the way.
Question 2Q: Single agent vs. multi-agent for a document-processing workflow: what does each side have going for it?Think it through, then reveal
A: For one agent (or a workflow): documents flow through the same steps, hand-offs lose context, multi-agent multiplies tokens, and one trace is far easier to debug. For several agents: documents are independent and can be processed in parallel, each document or section may overflow a single context, and specialists (tables, legal clauses, figures) benefit from focused instructions and tools. A common answer is a workflow that fans out per document to a focused worker, which is orchestrator-workers rather than free-form agents talking to each other.
Question 3Q: A router's classifier sometimes returns labels that aren't in its list. How should the code handle that?Think it through, then reveal
A: Normalize (case, whitespace), accept only known labels, and send
everything else to a safe fallback. Log the misses and add them to the
classifier's evaluation set. Consider structured output with an enum so
the model can only produce valid labels.
Question 4Q: Why checkpoint after every step of a long workflow?Think it through, then reveal
A: So a crash costs one step, not the run. Resuming reuses completed work (no repeated model calls, no double-posted invoices when combined with idempotency keys) and avoids getting a different answer on the rerun.
Primary sources
The papers behind this lesson
It established the reason-act-observe loop that the "agent loop" rung of the spectrum is built on.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Anthropic, Building effective agents (the named patterns): https://www.anthropic.com/engineering/building-effective-agents
- LangGraph docs: https://langchain-ai.github.io/langgraph/
- OpenAI Agents SDK: https://openai.github.io/openai-agents-python/
- CrewAI docs: https://docs.crewai.com/
- AutoGen: https://microsoft.github.io/autogen/
- Temporal, durable execution: https://docs.temporal.io/
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.