The lesson in one minute
What you'll be able to explain
- The model writes a request (
tool_use). Your code validates it, runs it, and returns atool_result. - Validate shape (JSON Schema) and meaning (business rules). Return every problem, each with its fix.
- Descriptions are prompts: say what the tool does, when to use it, and when not to.
- Prefer fewer, higher-level tools: punishes long chains of calls.
- Past a few dozen tools, load only the relevant ones per request.
- Writes get idempotency keys; irreversible or high-value actions get human approval; credentials get the narrowest scopes.
Level 1
The practitioner's guide
In one sentence
A tool is a function you let the model ask for, and tool design is everything that decides whether the request it fills in is the right one, has correct arguments, and is safe to carry out: the definition the model reads, the checks your code runs, and the gates in front of the real system.
When you need it
The moment a model's output does something rather than says something: looks up a record, refunds a payment, sends an email. A chatbot that only answers from its weights has no tools and needs none of this. A one-off script with one read-only tool needs the definition and little else. Everything in this lesson becomes necessary as soon as a tool writes, costs money or can be called by a model that was misled. The tell: if a wrong tool call would be visible to a customer, an auditor or a bank, you need the checks. Two of this lesson's measurements show how much the design decides. The same three tools, described vaguely ("search stuff"), get the right tool for 2 of 6 requests; described precisely (what it does, when to use it, when not to), 6 of 6. And a job that takes five tool calls in a row, each 97% likely to be right, succeeds 85.9% of the time, so about one run in seven fails; as one high-level call it succeeds 97% of the time.
Your options
Six layers of protection between the model's request and the real system, from the cheapest to the most certain. They stack; a production registry has all of them.
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| A schema in the definition | A JSON Schema tells the model which fields exist, their types and which are required | Nothing by itself; the model does its best to match | A few tokens per tool per call | Your tool definition |
| Strict mode | The provider constrains the model's output, token by token, to your schema | Arguments in exactly the schema's shape and a tool name that exists (Claude's strict tool use; OpenAI's strict mode) | A schema written to the provider's rules (additionalProperties: false, required fields); the compiled grammar is cached |
The model server |
| Validation in your registry | Checks shape and then business rules, and returns every problem with what a correct value looks like | A wrong call never reaches the system, and the model can fix all its mistakes in one retry (four at once in the lesson's example) | A validator and the rules; error messages worth writing | Your code |
| Descriptions and consolidation | Precise descriptions, few parameters, enums for closed sets, and one high-level tool in place of a chain |
The right tool chosen far more often (6 of 6 against 2 of 6 here), and fewer chances to fail in a row | Writing, and an evaluation set to measure selection | Your tool definitions |
| Loading only the relevant tools | Retrieval over the catalogue picks the few tools a request needs (5 of 30 in the lesson), or the provider searches it on demand | Selection accuracy holds as the catalogue grows; far fewer definition tokens per call | One embedding or search per request; the top few most-used tools kept always loaded | Your code, or the model server (Claude's tool search) |
| Safety gates | Scopes per agent, a dry run that describes instead of acting, human approval over a threshold, an idempotency key on every write | Least privilege, rehearsals, a person's sign-off, and a retry that cannot repeat a refund | A credential model, an approver in the loop, a store of keys and results | Your code, directly in front of the real system |
How to choose
Decide by what the tool can do, not by how clever the model is.
- Read-only lookups: a schema, a good description, and validation that returns actionable errors. Strict mode if the provider offers it; it removes a class of retries for free.
- Anything that writes: all of the above, plus an idempotency key on the
call and business-rule checks before it runs.
C-999fits the pattern^C-\d+$and is still a customer that does not exist. - Anything irreversible or over a money threshold: an approval rule, and a declined message that tells the model not to retry.
- Any agent that reads untrusted text (email, web pages, documents): the narrowest scopes you can give it, so a tricked model cannot refund, not merely should not.
- A catalogue past a few dozen tools: load per request, or route to a sub-agent that holds only its own tools. Claude's docs put the accuracy drop past 30 to 50 tools and recommend search from 10 tools up.
- Whatever you pick, prefer fewer, higher-level tools. A chain of five calls compounds five chances to go wrong; the sequencing belongs in tested code, not in the model.
What it costs
Tokens: every definition is re-sent on every call, so a
long catalogue is a standing charge. Claude's tool-search docs put a
typical multi-server setup (GitHub, Slack, Sentry, Grafana, Splunk) at about
55,000 tokens of definitions before any work is done, and on-demand loading
cuts that by over 85%. Tool results cost tokens too: Anthropic's tool-writing
guide reports that a concise response format used about a third of the
tokens of a detailed one, and that Claude Code caps a tool response at
25,000 tokens by default. Latency: validation is microseconds; an approval
gate is however long a person takes, so the call parks as
awaiting_approval rather than blocking. Quality: descriptions are the
cheapest lever, and the lesson's 2-of-6 to 6-of-6 swing cost only words.
Effort: the registry in this lesson, with every gate, is a few hundred lines,
and the idempotency store can be a dictionary until it needs to survive a
restart.
What breaks
- Schema-valid, false. Strict mode and schemas prove shape, not truth. Keep the business-rule check, and return an error that says how to find the right id.
- "Invalid input." A bare error makes the model guess. List every problem with the expected format ("date as YYYY-MM-DD"); the model fixes them all in one retry.
- Vague descriptions. A model that cannot tell the tools apart picks the first one every time; in the lesson it got exactly the two requests that happened to belong to that tool.
- Too many tools. Selection accuracy falls and definition tokens climb together. Load per request.
- A retry that pays twice. A refund that times out after the bank processed it gets retried. Without a key the customer gets $40; with the same key the registry returns the stored result and refunds once. Stripe's API keeps a key's first result and returns it on every repeat.
- A persistent model after a decline. "Declined" alone invites another attempt. Say "do not retry, tell the user".
- One admin credential. If the agent is tricked, it can do anything the credential allows. Scope per tool and per agent.
In the wild
Claude's Messages API accepts strict: true on a tool and
guarantees the arguments match the schema and the name is valid; its tool
search tool defers tool definitions and loads them when the model searches
for them; OpenAI's function calling has a strict mode its guide recommends
always enabling. Anthropic's Writing effective tools for agents is the
design reference behind sections 3 and 4 here: a few thoughtful tools over
one per endpoint, consistent namespacing (asana_search, jira_search),
responses that return meaning rather than identifiers, and evaluations
built from dozens of real prompt and response pairs. Stripe's idempotent
requests are the canonical form of the idempotency key: a client-chosen
key, the first result saved and replayed, keys pruned after 24 hours.
Toolformer (Schick et al., 2023) is the paper that showed a model can learn
when and how to call an API, which is why today's models emit tool calls
at all.
Go deeper
Level 2 writes the tool call out message by message, walks a call through every diamond of the registry in the order they are checked, measures the description experiment and the compounding formula, builds tool retrieval from cosine similarity, and replays the lost-reply refund with and without a key. If you only needed to choose, you are done.
Level 2
How it works, from scratch
A language model can only produce text. A tool is how that text turns into an action: looking up a record, refunding a payment, sending an email. This lesson builds a tool registry from scratch and shows the six things that decide whether tools work in production: how a call really happens, validation, descriptions, granularity, how many tools you expose, and safety.
Chapter 1
What a tool call really is
Everyday picture You hire a new assistant who is brilliant but may not touch anything. When they need something done, they fill in a request form ("refund customer C-100, $20") and hand it to you. You decide whether to carry it out, and you hand back a note saying what happened. The assistant is the model. The form is a tool call. You are your code.
Tiny worked example Here's the whole exchange for "what's 2 + 3?" with one tool, written as the actual messages:
you -> model tools: [{"name": "add", "description": "Add two integers.",
"input_schema": {"type": "object",
"properties": {"a": {"type": "integer"}, "b": {"type": "integer"}},
"required": ["a", "b"]}}]
user: "what's 2 + 3?"
model -> you tool_use {"id": "toolu_1", "name": "add", "input": {"a": 2, "b": 3}}
stop_reason: "tool_use" <- "please run this for me"
you (validate the input, run add(2, 3) -> 5)
you -> model user: tool_result {"tool_use_id": "toolu_1", "content": "5"}
model -> you "2 + 3 = 5." stop_reason: "end_turn"
input_schema is a JSON Schema, a JSON document that describes the
shape other JSON must have: which fields exist, what type each is, and
which are required. It's the "form template" the model fills in.
Figure 1 · Diagram
sequenceDiagram participant M as Model participant C as Your code participant T as The real system C->>M: tool definitions (name, description, JSON Schema) + user message M-->>C: tool_use: name + arguments (a filled-in form) C->>C: check the form (schema, business rules, permissions, approval) C->>T: carry it out T-->>C: result C->>M: tool_result (matched by tool_use_id) M-->>C: answer in words
The code ToolRegistry.register stores a Python function with its
name, description and schema; ToolRegistry.definitions() produces the
list you send to the model; ToolRegistry.call() runs a call through every
check.
Why it matters The separation is the security model. A model can be wrong or tricked. Your code is where you enforce what's allowed.
Chapter 2
Validate every call before running it
Everyday picture A bank clerk checks a withdrawal slip twice. Is it filled in properly: every box, a real date, a positive amount? Then does it make sense: does this account exist, does it have the money? Only then do they open the drawer.
Tiny worked example A payment tool takes an amount, a currency
(EUR, GBP or USD) and an optional pay_on date. The model sends
{"ammount": 5, "currency": "usd", "pay_on": "next friday"}. The validator
returns every problem at once, each one saying what a correct value looks
like:
$: missing required field 'amount'
$: unexpected field 'ammount' (allowed: amount, currency, pay_on)
$.currency: must be one of ['EUR', 'GBP', 'USD'], got 'usd'
$.pay_on: 'next friday' does not match the required format. Expected: date as YYYY-MM-DD
Compare that to invalid input. With the first, the model fixes all four
mistakes in one retry. With the second, it guesses.
Figure 2 · Diagram
flowchart TD
A[tool_use from the model] --> U{Tool exists?}
U -->|no| E1[unknown_tool:<br/>list the real tools]
U -->|yes| P{Caller has the<br/>required scope?}
P -->|no| E2[denied]
P -->|yes| S{Matches the<br/>JSON Schema?}
S -->|no| E3[invalid:<br/>every problem, with the fix]
S -->|yes| B{Business rules pass?<br/>e.g. customer exists}
B -->|no| E4[rejected]
B -->|yes| D{Dry run?}
D -->|yes| E5[dry_run:<br/>describe, change nothing]
D -->|no| H{Needs human<br/>approval?}
H -->|yes, none given| E6[awaiting_approval]
H -->|declined| E7[declined]
H -->|no, or approved| I{Idempotency key<br/>seen before?}
I -->|yes| E8[duplicate:<br/>return first result]
I -->|no| X[run the tool: ok]
status in
ToolOutcome, and each returns text the model can act on. The order is
deliberate: permissions come before anything that reveals how the tool works, and
cheap checks come before expensive ones. Approval and idempotency sit last,
right next to the action they protect.Why it matters "Structured output" and strict mode guarantee the
shape of the arguments, not that they're true. "C-999" matches the
pattern ^C-\d+$ perfectly and is still a customer that doesn't exist.
Schema checks and business-rule checks are both required.
With the Claude API you can add "strict": true to a tool definition
(definitions(strict=True) here). The API then guarantees the model's
arguments validate against the schema. You still need the business-rule
checks.
In code: validate checks a value against a JSON Schema subset and
returns every problem with its path. ToolRegistry.call walks the diamonds
above in order and wraps the result in a ToolOutcome, whose
ToolOutcome.is_error says whether the model should treat it as a failure.
Chapter 3
Descriptions are prompts
Everyday picture A wall of drawers labelled "stuff", "things" and "items". Even a careful person opens the wrong one. Relabel them "Policies (PTO, travel, VPN). Not customers", and nobody hesitates.
Tiny worked example A stand-in model (pick_tool) chooses the tool
whose name and description share the most words with the request. For
"billing status for customer 1042", the vague descriptions share zero words
with every tool (a three-way tie, so it guesses the first). The precise
lookup_record description shares "billing", "status" and "customer" and
wins.
Figure 3 · Drawn from the lesson's code
Vague descriptions pick the right tool for 2 of 6 requests; precise descriptions of the same tools get all 6 right
Why it matters Real models are far better readers than word overlap,
but they choose from exactly the same text. Precise descriptions, a few
parameters, enums instead of free text, and error messages that explain
the fix remove a large share of agent errors.
In code: selection_accuracy runs pick_tool over the labelled
requests and counts the correct picks, which is what the figure plots for
VAGUE_TOOLS and PRECISE_TOOLS.
Chapter 4
Fewer, higher-level tools
Everyday picture Asking an assistant to "book my trip to Denver" versus dictating five separate forms (find flight, hold seat, find hotel, reserve room, add to calendar). Every hand-off is another chance to drop something.
Figure 5 · Diagram
flowchart LR
subgraph Low["Five low-level calls"]
a1[get_customer] --> a2[get_price_list] --> a3[compute_tax] --> a4[render_pdf] --> a5[send_email]
end
subgraph High["One high-level call"]
b1[create_invoice]
end
The chance that a chain of calls all succeed:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning |
|---|---|
| probability that one call is chosen and filled in correctly | |
| number of calls in the chain |
In words: multiply the per-call success rate by itself once per call.
On the example: with and , , so about 1 run in 7 fails. With one high-level call it's .
In Python:
p = 0.97
# five calls that must all succeed
round(p ** 5, 3) # → 0.859
# about 1 run in 7 fails
round(1 - p ** 5, 2) # → 0.14
# one high-level call
p ** 1 # → 0.97
Figure 4 · Drawn from the lesson's code
Success falls with every extra call: at 99% per call twenty calls succeed about 82% of the time, at 90% ten calls barely a third
In code: chain_success evaluates .
Chapter 5
Too many tools: load only the relevant ones
Everyday picture A warehouse versus a toolbox. You don't send a plumber into a warehouse of 500 tools; you hand them the five they need for today's job.
Tiny worked example A catalogue of 30 tools. For "I forgot my password and I'm
locked out", select_tools embeds the request, compares it to every tool
description, and sends the model only the top 5, with reset_password
first.
Figure 7 · Diagram
flowchart LR R[User request] --> E[Embed request] C[(Tool catalogue<br/>30 descriptions,<br/>embedded once)] --> S E --> S[Cosine similarity<br/>to every tool] S --> K[Top 5 tools] K --> M[Model sees<br/>only these 5]
Figure 6 · Drawn from the lesson's code
Each request is similar to only a small cluster of related tools and dim against the rest of the catalogue
Why it matters Selection accuracy drops as the tool list grows past a few dozen, and every definition costs input tokens on every call. The alternatives are routing the request to a sub-agent that holds only the relevant tools, or using a provider's built-in tool search.
Chapter 6
Safety: idempotency, dry runs, approval, least privilege
Idempotency. Everyday picture: pressing a lift button twice still takes you to the floor once. An operation is idempotent when doing it twice has the same effect as doing it once. Worked example: a refund call times out after the bank processed it, so the agent retries.
Figure 8 · Diagram
sequenceDiagram participant A as Agent participant R as Registry participant B as Bank A->>R: refund C-100 $20 (key req-1) R->>B: refund B-->>R: done R--xA: network timeout (reply lost) A->>R: retry: refund C-100 $20 (key req-1) R-->>A: duplicate: "refunded 20.00 to C-100" (no second refund)
Dry run. A rehearsal: run every check, then describe the action
instead of doing it. Useful for previews and for shadow-mode rollouts
(primer.agents.deployment).
Human approval. Everyday picture: a manager signs off on spending over
a limit. Irreversible actions, or ones over a threshold such as refunds over
$500, park as awaiting_approval until a person decides.
Figure 9 · Diagram
sequenceDiagram
participant M as Model
participant R as Registry
participant H as Human approver
M->>R: refund C-100 $900
R->>H: approve refund_payment {amount: 900}?
alt approved
H-->>R: yes
R-->>M: ok: refunded
else declined
H-->>R: no
R-->>M: declined: do not retry, tell the user
end
Least privilege. Everyday picture: a valet key starts the car but
won't open the boot. Give each agent credentials (scopes, named
permissions such as payments:write) for its job only. An agent that reads
tickets holds tickets:read, so even if a malicious email tricks it, it
cannot issue refunds.
In code: a Tool carries its safety settings: the scopes it needs,
whether it reads, writes or acts irreversibly, and an optional approval
rule, which Tool.requires_approval combines. ToolRegistry.call enforces
them, given the caller's credentials, an idempotency key, a dry-run flag and
an approver.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: How do you design tools so the model picks the right one with correct arguments?Think it through, then reveal
A: Give each tool one clear job and a description that says what it does,
when to use it and when not to. Keep parameters few, use enums for closed
sets, and put the format in the schema description ("date as YYYY-MM-DD").
Prefer one high-level tool over several low-level ones. Validate every call
and return actionable errors so the model can self-correct. Measure tool
selection accuracy in evals, and load tools dynamically when there are many.
Question 2Q: The schema says customer_id must match ^C-\d+$. Is that enough validation?Think it through, then reveal
A: No. The schema checks shape, not truth. C-999 is well-formed and may not
exist. Add a business-rule check before acting, and return an error that
tells the model how to find the right id.
Question 3Q: A refund call timed out. Should the agent retry?Think it through, then reveal
A: Only if the call is idempotent. Send an idempotency key with every write, and the retry returns the original result instead of refunding twice.
Question 4Q: Why not give the agent one admin credential for everything?Think it through, then reveal
A: Blast radius. If the agent is tricked (for example by prompt injection in a document it reads), it can do anything the credential allows. Scope credentials per tool and per agent, and put irreversible actions behind human approval.
Primary sources
The papers behind this lesson
It showed a language model can learn when to call an API, which one and with what arguments, by keeping only the self-generated calls that made its predictions better, which is the idea behind tool calling being trained into today's models.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Anthropic, Writing effective tools for agents: https://www.anthropic.com/engineering/writing-tools-for-agents
- Anthropic, Building effective agents (appendix on prompt-engineering tools): https://www.anthropic.com/engineering/building-effective-agents
- Claude tool use overview: https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
- Implementing tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
- Understanding JSON Schema: https://json-schema.org/understanding-json-schema
- Stripe, idempotent requests (the canonical explanation): https://docs.stripe.com/api/idempotent_requests
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.