rumblr Work in progressWIP

● The AI Primer · Lesson 39 · Part 2: building systems people rely on

Orchestration

how much autonomy to give the model

You'll be able to explain Workflows vs. agents, and the named patterns

Members · open during launch 21 min11 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Autonomy spectrum: fixed workflow → router → agent loop → multi-agent. Use the least that works.
  2. Named patterns: chaining (with gates), routing (with a fallback), parallelization (sectioning, voting), orchestrator-workers, evaluator-optimizer.
  3. Multi-agent buys focused context and parallelism, and costs tokens, hand-off losses and debuggability.
  4. Production shape: an explicit state machine, the model consulted only at specific nodes, checkpoints after each step.

Level 1

The practitioner's guide

In one sentence

Orchestration is the decision of how much of a system's control flow the model gets to decide, from none (code runs fixed steps and asks the model to fill each one) to all of it (a team of models steering each other), and the named patterns in between.

When you need it

You face this decision the moment a task needs more than one model call: a document pipeline, a support desk, a research assistant, anything that must look something up and then act. The tell that you have chosen wrongly in one direction is a fixed pipeline that keeps growing special cases for requests it didn't anticipate; the tell in the other is an agent that spends a minute and a thousand tokens on a request a two-step script answered correctly every time. Anthropic's Building effective agents, the post these pattern names come from, draws the line this way: workflows are model calls "orchestrated through predefined code paths"; agents are systems where the model "dynamically directs its own processes and tool usage". It also gives the rule: find the simplest solution possible, and add complexity only when needed. You don't need orchestration at all when one call with retrieval and a few examples in the prompt answers the question; that post says many applications stop there.

This lesson measures the cost of each step up on one question ("How many PTO days do I get, and what is the meal per-diem?"). A fixed workflow answers it in 2 calls and 93 tokens; a router in 2 calls and 105 tokens; an agent loop in 2 calls and about 990 tokens, ten times the workflow, because it re-sends its tool definitions and results on every call; a supervisor with two specialists in 4 calls and about 600 tokens. Same answer, four prices.

Your options

Seven designs, from the least autonomy to the most:

Option What it does What it guarantees What it costs Where it lives
Fixed workflow (prompt chaining with gates) Code runs steps in order; the model fills each one; a code check between steps stops a bad result The same steps every time; a failed gate stops before the next call is paid for or misled 2 calls and 93 tokens on the lesson's question; a gate per hand-off to write Your code
Router One cheap classifier call picks a handler; unknown labels go to a fallback Exactly one path per request, and never a crash on a label the model invented One extra call (105 tokens here) and a fallback handler Your code, with one model decision
Parallelization (sectioning, voting) Independent pieces run at once, or several judges answer the same prompt and the majority wins Wall-clock time of the slowest branch; three 80%-accurate judges make an 89.6% panel if they err independently n times the tokens for n branches Your code (a thread pool)
Evaluator-optimizer One call drafts, another critiques, until it passes or the round limit hits A checked result, or a flagged best effort at the limit Up to two calls per round (2 rounds in the lesson's example) Your code
Agent loop The model picks tools and decides when it is done Handles requests you did not anticipate About 990 tokens for the lesson's question, ten times the workflow; less predictability The model decides; your loop enforces the budget (primer.agents.agent_loop)
Orchestrator-workers, multi-agent A lead model splits the work, specialists each get only their brief, a synthesizer combines Focused context per specialist and parallel work; subtasks decided at run time 4 calls and about 600 tokens here; in Anthropic's production research system, about 15 times the tokens of a chat Several model roles, coordinated by your code
Explicit state machine with checkpoints Code owns every state and transition; the model is consulted at a few named nodes; the run is saved after each step Auditable, testable transitions and a crash that costs one step, not the run A state model to design and a store for checkpoints Your code, or a durable-execution engine

How to choose

Start from whether the steps vary between requests, or only what happens inside each step.

  • Steps known in advance, variation inside them (invoice intake: classify, extract, validate, approve, post): a fixed workflow or state machine, with the model at the judgement nodes only.
  • A few kinds of request, each with its own handling: a router in front of the workflows, with a fallback that is safe rather than clever.
  • One hard judgement call: voting, but vary the prompt, model or evidence, because copies of the same model tend to make the same mistake.
  • Quality that a checker can state as criteria (a citation present, a test passing): evaluator-optimizer, with a round limit.
  • The path depends on what is discovered along the way: an agent loop.
  • Work that is separable and too big for one context (independent research threads, per-document processing): orchestrator-workers. Not for tasks where every part needs the same context or depends on the others; the same Anthropic report found most coding tasks fall on that side.
  • Whatever you pick, code owns the transitions you need to audit, and the model decides only where judgement is needed.

What it costs

Tokens, first: each rung up the spectrum multiplies them, from 93 to 105 to about 990 to about 600 on one toy question, and Anthropic reports agents using about 4 times the tokens of a chat and multi-agent systems about 15 times. Latency: a chain is the sum of its calls, a parallel fan-out is its slowest branch, and every hand-off in a multi-agent design adds a call. Quality: multi-agent can win (that report measured a 90.2% improvement over a single agent on an internal research evaluation) when the work parallelizes, and loses context at every hand-off when it does not. Effort: every pattern in this lesson is under 100 lines of plain Python; a framework adds tracing, persistence and approval hooks, and hidden assumptions. Crashes: without checkpoints, a rerun after a crash repeats every model call (4 instead of 2 in the lesson's invoice run) and risks a different answer the second time.

What breaks

  • A router that trusts the label. The classifier returns free text ("Billing-ish, maybe refunds?"). Normalize it, accept only known labels, send the rest to a fallback, and log the misses.
  • A gate that isn't there. Without a check between chain steps, a one-line outline becomes a paid-for, confident draft of nothing.
  • Correlated voters. Five copies of one model with one prompt agree on the same error; the majority formula assumes independence.
  • An evaluator that is never satisfied. No round limit, no exit. Put the limit in, and flag what came out at the limit.
  • Lost hand-offs. A specialist knows only what the supervisor wrote in its brief. Whatever the supervisor forgot is gone. Test that each specialist sees only its own task, and that a delegation to a specialist that doesn't exist is reported, not crashed on.
  • No checkpoint. A crash on the last step restarts from the first, doubling cost and, combined with a non-idempotent side effect, posting an invoice twice.

In the wild

The pattern names come from Anthropic's Building effective agents, which also advises starting with the model API directly before reaching for a framework. LangGraph models a workflow as a graph with explicit state and checkpoints, which is the state-machine row above. The Claude Agent SDK and the OpenAI Agents SDK ship vendor-built agent loops, tools and hand-offs. CrewAI and AutoGen are built around multi-agent teams. Temporal provides durable execution for any code: each completed step's result is stored and a crashed run resumes from there. Anthropic's own research feature is an orchestrator-workers system with a lead model and parallel subagents, and its engineering report is the source of the token multipliers above.

Go deeper

Level 2 builds every pattern in a few dozen lines each and measures it: the four-rung cost table, a chain with a gate, a router with a fallback, voting with the formula for why independent judges help, a supervisor whose specialists see only their brief, and an invoice state machine that survives a crash. If you only needed to choose, you are done.

Level 2

How it works, from scratch

"Agent" covers everything from one model call inside ordinary code to a team of models steering each other. This lesson lays out that range, builds each named pattern in a few dozen lines, and measures what each step up costs. The rule that falls out is simple: use the least autonomy that solves the problem, and say so out loud.

Chapter 1

The autonomy spectrum

Everyday picture Four ways to run an office. An assembly line (fixed workflow): the steps never change and people do one task at each station. A receptionist (router) decides which department you go to. A personal assistant (agent loop) decides which errands to run and when the job's done. A team with a manager (multi-agent) splits the work among specialists.

Figure 2 · Diagram

Reading it: left to right, the model decides more and your code decides less. Each arrow buys flexibility (the system can handle requests you didn't anticipate) and costs predictability, tokens and debuggability. The skill is stopping at the leftmost box that handles your real requests.

Tiny worked example One question, "How many PTO days do I get, and what is the meal per-diem when travelling?", answered by each design (autonomy_costs()):

design model calls why
fixed workflow 2 rewrite as a search, then answer
router 2 classify, then the chosen handler answers
agent loop 2 one turn with two parallel tool calls, one answer turn
multi-agent 4 plan, one call per specialist (2), final answer

Figure 1 · Drawn from the lesson's code

fixed workflow router agent loop multi-agent 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 model calls Calls for the same question fixed workflow router agent loop multi-agent 0 200 400 600 800 1000 tokens (input + output) Tokens for the same question

Workflow, router and agent loop each make 2 calls and multi-agent 4, but the agent loop uses about 990 tokens, ten times the workflow's 93

Reading it: the left panel counts model calls and the right counts tokens. Calls barely move until multi-agent doubles them. Tokens tell a sharper story. The agent loop re-sends its tool definitions and tool results on every call, so it uses many times the tokens of the fixed workflow for the same answer. The supervisor pays for planning and hand-offs. None of that is waste if you need the flexibility, and all of it is waste if you don't.

Chapter 2

The named patterns

These names (from Anthropic's Building effective agents) let you describe a design in one word.

In code: every pattern below is built from ask, one model call with one user message that returns the reply's text.

Prompt chaining

Everyday picture A relay race with an inspector at each hand-off: if the baton is dropped, the next runner doesn't start.

Tiny worked example Step 1 writes an outline, and a code gate checks it has at least three points. The outline "- scope" fails the gate, so the chain stops after one call and the draft step never runs.

Figure 4 · Diagram

Reading it: two model calls in a fixed order, with plain code in the middle. The gate is where you catch a bad intermediate result before paying for, and being misled by, the next step.

In code: run_chain runs a list of ChainSteps in order, feeding each output into the next prompt and stopping at the first gate that reports a problem; ChainResult records the outputs and where and why it stopped.

Routing

Everyday picture The receptionist listens for five seconds and sends you to billing, tech support or the front desk.

Tiny worked example The classifier replies " Technical\n", which is normalized to technical and routed. It replies "Billing-ish, maybe refunds?", which is not a known label, so the request goes to the general fallback instead of crashing.

Figure 5 · Diagram

Reading it: one cheap model call decides the path, then specialised handling takes over. The "anything else" arrow is essential, because the label is free text from a model and code must never assume it's one of the keys.

In code: route makes the one classifier call, normalizes the label, swaps anything unknown for the fallback, and hands the request to that label's handler.

Parallelization: sectioning and voting

Everyday picture Sectioning: several cooks each make one dish at the same time. Voting: a panel of judges, where the majority decides.

Tiny worked example Three reviewers judge a SQL snippet: two say "vulnerable" and one says "safe", so the verdict is ("vulnerable", {"vulnerable": 2, "safe": 1}).

Figure 6 · Diagram

Reading it: the three calls run at the same time (wall-clock time is the slowest one), and plain code counts the answers. Sectioning looks the same except each branch gets a different sub-task, such as a guardrail check running beside the answer.

Why voting helps, when voters err independently:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
number of voters (odd, so there are no ties)
chance a single voter is right
number of voters who are right
ways to choose which of the are right

In words: add up the chances of every outcome where more than half the voters are right.

On the example: , : . Three 80% judges make an 89.6% panel.

In Python:

from math import comb
n, p = 3, 0.8
# the chance exactly k voters are right ...
P = sum(comb(n, k) * p ** k * (1 - p) ** (n - k)
        # ... for every k from ⌊n/2⌋ + 1 to n
        for k in range(n // 2 + 1, n + 1))
round(P, 3)  # → 0.896

Figure 3 · Drawn from the lesson's code

2 4 6 8 10 12 14 number of independent voters (odd) 0.5 0.6 0.7 0.8 0.9 1.0 P(majority is right) Voting: many noisy judges beat one (if their errors are independent) each voter right 60% each voter right 70% each voter right 80% each voter right 90%

Majority accuracy rises with more independent voters: at 70% per voter, 5 voters reach 84% and 15 reach 95%; at 60% per voter the climb is slow

Reading it: each curve is a per-voter accuracy, and the x-axis adds voters. Every curve rises: more independent judges, better majority. The catch is the word independent. Five copies of the same model with the same prompt tend to make the same mistake, and then voting buys little. Vary the prompt, the model or the evidence.

In code: run_sections is sectioning: it runs independent pieces of work on a thread pool and collects every result. vote is voting: it asks each reviewer the same prompt through run_sections and returns the majority label with its tally, and majority_accuracy evaluates the formula above.

Orchestrator-workers

Everyday picture A project lead reads the brief, splits it into sections, hands each to a writer, then edits the pieces into one report.

Tiny worked example "What's the meal per-diem, when are receipts due, and does PTO roll over?" The orchestrator returns three subtasks. Three workers each find one document (fin-002, fin-004, hr-001), and the synthesizer writes one answer citing all three.

Figure 7 · Diagram

Reading it: it looks like sectioning, but the subtasks aren't known in advance. The orchestrator decides them from the request, which is what makes it suited to open-ended requests.

In code: orchestrate asks the orchestrator for subtasks, runs the workers on them in parallel with run_sections, and asks the synthesizer for one answer, returning an OrchestratorResult. policy_worker is the worker in the example: it finds the one current policy document for a subtask.

Evaluator-optimizer

Everyday picture A writer and an editor. The draft goes back and forth until the editor signs off, or the deadline hits.

Tiny worked example Draft 1, "PTO rolls over.", gets the feedback "Missing a citation". Draft 2, "Up to 5 unused PTO days roll over [hr-001].", passes. Result: 2 rounds, passed.

Figure 8 · Diagram

Reading it: the loop only exits on PASS or on the round limit. Without the limit, a critic that's never satisfied loops forever. Clear, checkable criteria make this pattern work.

In code: evaluate_optimize alternates generator drafts and evaluator verdicts, feeding each critique back, and returns the last draft, the rounds used and whether it passed.

Chapter 3

Multi-agent: a supervisor and specialists

Everyday picture A manager who never does the work, only assigns it and assembles the results. Each specialist gets a short, focused brief.

Tiny worked example The supervisor sends "PTO days per year?" to hr and "Meal per-diem?" to finance. The HR specialist never sees the finance question, and vice versa (tested). A delegation to a non-existent legal specialist is reported, not crashed on.

Figure 9 · Diagram

Reading it: the supervisor's messages to each specialist are the hand-off, and they're all each specialist knows. That's both the benefit (small, focused context) and the risk (whatever the supervisor forgets to write down is lost). Peer designs, where agents hand control directly to each other, have the same hand-off problem without a central coordinator.

When it's worth it. Use several agents when subtasks are genuinely separable, such as independent research threads or per-document work that would overflow one context. The costs are real: more tokens, context lost at every hand-off, and much harder debugging.

In code: supervise gets a delegation plan from the supervisor, sends each specialist only its own task (reporting unknown names instead of crashing), then asks the supervisor for the final answer; SupervisorResult keeps the delegations and every specialist's answer.

Chapter 4

An explicit state machine, with the model at specific nodes

Everyday picture A board game. The squares and the rules for moving between them are fixed. On a few special squares you draw a card (ask the model). Everywhere else, the rules decide.

Tiny worked example Invoice intake. The model is asked twice: "is this an invoice?" and "extract the number, vendor and amount". Code does everything else: validation, the approval rule (over 5,000 waits for a person), and posting. The runs:

INVOICE INV-2041 ... 1200.00   -> RECEIVED > CLASSIFIED > EXTRACTED > VALIDATED > POSTED > DONE
INVOICE INV-2042 ... 18000.00  -> RECEIVED > CLASSIFIED > EXTRACTED > VALIDATED > AWAITING_APPROVAL
Acme newsletter                -> RECEIVED > REJECTED

Figure 11 · Diagram

Reading it: every box is a state and every arrow a transition, and only two arrows are labelled "LLM". The rest are rules you can read, test and audit. When something goes wrong, you know exactly which state it was in and why it moved.

In code: InvoiceWorkflow.run walks the states, calling the model only at RECEIVED and CLASSIFIED; validate_invoice is the code check between EXTRACTED and VALIDATED; InvoiceWorkflow.approve is the human step that releases a paused invoice. WorkflowRun holds the current state, its history and the extracted data.

Checkpoints and durable execution. After every transition the run is saved to disk. Durable execution means a workflow whose progress survives crashes, because each completed step's result is stored and the run resumes from there. Engines like Temporal do this at scale.

Figure 10 · Drawn from the lesson's code

resume from checkpoint restart from scratch 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 model calls to finish one invoice Posting crashed once: what did it cost? 2 4

After one crash while posting, resuming from the checkpoint finishes in 2 model calls, while restarting from scratch takes 4

Reading it: the same crash happens in both bars. Resuming from the checkpoint finishes with the 2 model calls already made. Restarting from scratch asks the model both questions again, doubling cost, and, worse, risks a different answer the second time.

In code: InvoiceWorkflow saves a checkpoint file after every transition and loads it at the start of InvoiceWorkflow.run, so a rerun resumes. resume_comparison crashes the ledger once and counts the model calls each way.

Chapter 5

Frameworks vs. plain code

LangGraph models agents as graphs with explicit state and checkpoints. The Claude Agent SDK and OpenAI Agents SDK give you vendor-built agent loops, tools and hand-offs. CrewAI and AutoGen focus on multi-agent teams. Temporal provides durable execution for any code. A framework earns its place with standard patterns, built-in tracing, persistence and human-in-the-loop hooks. Plain code wins for simple flows, full control and fewer dependencies. Everything in this lesson is under 100 lines of plain Python, and knowing that is what lets you judge a framework.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: When would you use a fixed workflow instead of an agent? Give an example.Think it through, then reveal

A: When the steps are known in advance and the variation is inside each step, not in which steps happen. Invoice intake is the example: classify, extract, validate, approve, post. A workflow is cheaper, faster, testable and auditable. The model is used where judgement is needed (classification, extraction), and code owns the flow. Reach for an agent only when the path genuinely depends on what's discovered along the way.

Question 2Q: Single agent vs. multi-agent for a document-processing workflow: what does each side have going for it?Think it through, then reveal

A: For one agent (or a workflow): documents flow through the same steps, hand-offs lose context, multi-agent multiplies tokens, and one trace is far easier to debug. For several agents: documents are independent and can be processed in parallel, each document or section may overflow a single context, and specialists (tables, legal clauses, figures) benefit from focused instructions and tools. A common answer is a workflow that fans out per document to a focused worker, which is orchestrator-workers rather than free-form agents talking to each other.

Question 3Q: A router's classifier sometimes returns labels that aren't in its list. How should the code handle that?Think it through, then reveal

A: Normalize (case, whitespace), accept only known labels, and send everything else to a safe fallback. Log the misses and add them to the classifier's evaluation set. Consider structured output with an enum so the model can only produce valid labels.

Question 4Q: Why checkpoint after every step of a long workflow?Think it through, then reveal

A: So a crash costs one step, not the run. Resuming reuses completed work (no repeated model calls, no double-posted invoices when combined with idempotency keys) and avoids getting a different answer on the rerun.

Primary sources

The papers behind this lesson

Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022).

It established the reason-act-observe loop that the "agent loop" rung of the spectrum is built on.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Building effective agents (the named patterns): https://www.anthropic.com/engineering/building-effective-agents
  • LangGraph docs: https://langchain-ai.github.io/langgraph/
  • OpenAI Agents SDK: https://openai.github.io/openai-agents-python/
  • CrewAI docs: https://docs.crewai.com/
  • AutoGen: https://microsoft.github.io/autogen/
  • Temporal, durable execution: https://docs.temporal.io/

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.