At a glance
Key takeaways
- The model writes a request (
tool_use). Your code validates it, runs it, and returns atool_result. - Validate shape (JSON Schema) and meaning (business rules). Return every problem, each with its fix.
- Descriptions are prompts: say what the tool does, when to use it, and when not to.
- Prefer fewer, higher-level tools: punishes long chains of calls.
- Past a few dozen tools, load only the relevant ones per request.
- Writes get idempotency keys; irreversible or high-value actions get human approval; credentials get the narrowest scopes.
Level 2
How it works, from scratch
A language model can only produce text. A tool is how that text turns into an action: looking up a record, refunding a payment, sending an email. This lesson builds a tool registry from scratch and shows the six things that decide whether tools work in production: how a call really happens, validation, descriptions, granularity, how many tools you expose, and safety.
Chapter 1
What a tool call really is
Everyday picture You hire a new assistant who is brilliant but may not touch anything. When they need something done, they fill in a request form ("refund customer C-100, $20") and hand it to you. You decide whether to carry it out, and you hand back a note saying what happened. The assistant is the model. The form is a tool call. You are your code.
Tiny worked example Here's the whole exchange for "what's 2 + 3?" with one tool, written as the actual messages:
you -> model tools: [{"name": "add", "description": "Add two integers.",
"input_schema": {"type": "object",
"properties": {"a": {"type": "integer"}, "b": {"type": "integer"}},
"required": ["a", "b"]}}]
user: "what's 2 + 3?"
model -> you tool_use {"id": "toolu_1", "name": "add", "input": {"a": 2, "b": 3}}
stop_reason: "tool_use" <- "please run this for me"
you (validate the input, run add(2, 3) -> 5)
you -> model user: tool_result {"tool_use_id": "toolu_1", "content": "5"}
model -> you "2 + 3 = 5." stop_reason: "end_turn"
input_schema is a JSON Schema, a JSON document that describes the
shape other JSON must have: which fields exist, what type each is, and
which are required. It's the "form template" the model fills in.
Figure 1 · Diagram
sequenceDiagram participant M as Model participant C as Your code participant T as The real system C->>M: tool definitions (name, description, JSON Schema) + user message M-->>C: tool_use: name + arguments (a filled-in form) C->>C: check the form (schema, business rules, permissions, approval) C->>T: carry it out T-->>C: result C->>M: tool_result (matched by tool_use_id) M-->>C: answer in words
The code ToolRegistry.register stores a Python function with its
name, description and schema; ToolRegistry.definitions() produces the
list you send to the model; ToolRegistry.call() runs a call through every
check.
Why it matters The separation is the security model. A model can be wrong or tricked. Your code is where you enforce what's allowed.
Chapter 2
Validate every call before running it
Everyday picture A bank clerk checks a withdrawal slip twice. Is it filled in properly: every box, a real date, a positive amount? Then does it make sense: does this account exist, does it have the money? Only then do they open the drawer.
Tiny worked example A payment tool takes an amount, a currency
(EUR, GBP or USD) and an optional pay_on date. The model sends
{"ammount": 5, "currency": "usd", "pay_on": "next friday"}. The validator
returns every problem at once, each one saying what a correct value looks
like:
$: missing required field 'amount'
$: unexpected field 'ammount' (allowed: amount, currency, pay_on)
$.currency: must be one of ['EUR', 'GBP', 'USD'], got 'usd'
$.pay_on: 'next friday' does not match the required format. Expected: date as YYYY-MM-DD
Compare that to invalid input. With the first, the model fixes all four
mistakes in one retry. With the second, it guesses.
Figure 2 · Diagram
flowchart TD
A[tool_use from the model] --> U{Tool exists?}
U -->|no| E1[unknown_tool:<br/>list the real tools]
U -->|yes| P{Caller has the<br/>required scope?}
P -->|no| E2[denied]
P -->|yes| S{Matches the<br/>JSON Schema?}
S -->|no| E3[invalid:<br/>every problem, with the fix]
S -->|yes| B{Business rules pass?<br/>e.g. customer exists}
B -->|no| E4[rejected]
B -->|yes| D{Dry run?}
D -->|yes| E5[dry_run:<br/>describe, change nothing]
D -->|no| H{Needs human<br/>approval?}
H -->|yes, none given| E6[awaiting_approval]
H -->|declined| E7[declined]
H -->|no, or approved| I{Idempotency key<br/>seen before?}
I -->|yes| E8[duplicate:<br/>return first result]
I -->|no| X[run the tool: ok]
status in
ToolOutcome, and each returns text the model can act on. The order is
deliberate: permissions come before anything that reveals how the tool works, and
cheap checks come before expensive ones. Approval and idempotency sit last,
right next to the action they protect.Why it matters "Structured output" and strict mode guarantee the
shape of the arguments, not that they're true. "C-999" matches the
pattern ^C-\d+$ perfectly and is still a customer that doesn't exist.
Schema checks and business-rule checks are both required.
With the Claude API you can add "strict": true to a tool definition
(definitions(strict=True) here). The API then guarantees the model's
arguments validate against the schema. You still need the business-rule
checks.
In code: validate checks a value against a JSON Schema subset and
returns every problem with its path. ToolRegistry.call walks the diamonds
above in order and wraps the result in a ToolOutcome, whose
ToolOutcome.is_error says whether the model should treat it as a failure.
Chapter 3
Descriptions are prompts
Everyday picture A wall of drawers labelled "stuff", "things" and "items". Even a careful person opens the wrong one. Relabel them "Policies (PTO, travel, VPN). Not customers", and nobody hesitates.
Tiny worked example A stand-in model (pick_tool) chooses the tool
whose name and description share the most words with the request. For
"billing status for customer 1042", the vague descriptions share zero words
with every tool (a three-way tie, so it guesses the first). The precise
lookup_record description shares "billing", "status" and "customer" and
wins.
Figure 3 · Drawn from the lesson's code
Vague descriptions pick the right tool for 2 of 6 requests; precise descriptions of the same tools get all 6 right
Why it matters Real models are far better readers than word overlap,
but they choose from exactly the same text. Precise descriptions, a few
parameters, enums instead of free text, and error messages that explain
the fix remove a large share of agent errors.
In code: selection_accuracy runs pick_tool over the labelled
requests and counts the correct picks, which is what the figure plots for
VAGUE_TOOLS and PRECISE_TOOLS.
Chapter 4
Fewer, higher-level tools
Everyday picture Asking an assistant to "book my trip to Denver" versus dictating five separate forms (find flight, hold seat, find hotel, reserve room, add to calendar). Every hand-off is another chance to drop something.
Figure 4 · Diagram
flowchart LR
subgraph Low["Five low-level calls"]
a1[get_customer] --> a2[get_price_list] --> a3[compute_tax] --> a4[render_pdf] --> a5[send_email]
end
subgraph High["One high-level call"]
b1[create_invoice]
end
The chance that a chain of calls all succeed:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning |
|---|---|
| probability that one call is chosen and filled in correctly | |
| number of calls in the chain |
In words: multiply the per-call success rate by itself once per call.
On the example: with and , , so about 1 run in 7 fails. With one high-level call it's .
In Python:
p = 0.97
# five calls that must all succeed
round(p ** 5, 3) # → 0.859
# about 1 run in 7 fails
round(1 - p ** 5, 2) # → 0.14
# one high-level call
p ** 1 # → 0.97
Figure 5 · Drawn from the lesson's code
Success falls with every extra call: at 99% per call twenty calls succeed about 82% of the time, at 90% ten calls barely a third
In code: chain_success evaluates .
Chapter 5
Too many tools: load only the relevant ones
Everyday picture A warehouse versus a toolbox. You don't send a plumber into a warehouse of 500 tools; you hand them the five they need for today's job.
Tiny worked example A catalogue of 30 tools. For "I forgot my password and I'm
locked out", select_tools embeds the request, compares it to every tool
description, and sends the model only the top 5, with reset_password
first.
Figure 6 · Diagram
flowchart LR R[User request] --> E[Embed request] C[(Tool catalogue<br/>30 descriptions,<br/>embedded once)] --> S E --> S[Cosine similarity<br/>to every tool] S --> K[Top 5 tools] K --> M[Model sees<br/>only these 5]
Figure 7 · Drawn from the lesson's code
Each request is similar to only a small cluster of related tools and dim against the rest of the catalogue
Why it matters Selection accuracy drops as the tool list grows past a few dozen, and every definition costs input tokens on every call. The alternatives are routing the request to a sub-agent that holds only the relevant tools, or using a provider's built-in tool search.
Chapter 6
Safety: idempotency, dry runs, approval, least privilege
Idempotency. Everyday picture: pressing a lift button twice still takes you to the floor once. An operation is idempotent when doing it twice has the same effect as doing it once. Worked example: a refund call times out after the bank processed it, so the agent retries.
Figure 8 · Diagram
sequenceDiagram participant A as Agent participant R as Registry participant B as Bank A->>R: refund C-100 $20 (key req-1) R->>B: refund B-->>R: done R--xA: network timeout (reply lost) A->>R: retry: refund C-100 $20 (key req-1) R-->>A: duplicate: "refunded 20.00 to C-100" (no second refund)
Dry run. A rehearsal: run every check, then describe the action
instead of doing it. Useful for previews and for shadow-mode rollouts
(primer.agents.deployment).
Human approval. Everyday picture: a manager signs off on spending over
a limit. Irreversible actions, or ones over a threshold such as refunds over
$500, park as awaiting_approval until a person decides.
Figure 9 · Diagram
sequenceDiagram
participant M as Model
participant R as Registry
participant H as Human approver
M->>R: refund C-100 $900
R->>H: approve refund_payment {amount: 900}?
alt approved
H-->>R: yes
R-->>M: ok: refunded
else declined
H-->>R: no
R-->>M: declined: do not retry, tell the user
end
Least privilege. Everyday picture: a valet key starts the car but
won't open the boot. Give each agent credentials (scopes, named
permissions such as payments:write) for its job only. An agent that reads
tickets holds tickets:read, so even if a malicious email tricks it, it
cannot issue refunds.
In code: a Tool carries its safety settings: the scopes it needs,
whether it reads, writes or acts irreversibly, and an optional approval
rule, which Tool.requires_approval combines. ToolRegistry.call enforces
them, given the caller's credentials, an idempotency key, a dry-run flag and
an approver.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: How do you design tools so the model picks the right one with correct arguments?Think it through, then reveal
A: Give each tool one clear job and a description that says what it does,
when to use it and when not to. Keep parameters few, use enums for closed
sets, and put the format in the schema description ("date as YYYY-MM-DD").
Prefer one high-level tool over several low-level ones. Validate every call
and return actionable errors so the model can self-correct. Measure tool
selection accuracy in evals, and load tools dynamically when there are many.
Question 2Q: The schema says customer_id must match ^C-\d+$. Is that enough validation?Think it through, then reveal
A: No. The schema checks shape, not truth. C-999 is well-formed and may not
exist. Add a business-rule check before acting, and return an error that
tells the model how to find the right id.
Question 3Q: A refund call timed out. Should the agent retry?Think it through, then reveal
A: Only if the call is idempotent. Send an idempotency key with every write, and the retry returns the original result instead of refunding twice.
Question 4Q: Why not give the agent one admin credential for everything?Think it through, then reveal
A: Blast radius. If the agent is tricked (for example by prompt injection in a document it reads), it can do anything the credential allows. Scope credentials per tool and per agent, and put irreversible actions behind human approval.
Primary sources
The papers behind this lesson
It showed a language model can learn when to call an API, which one and with what arguments, by keeping only the self-generated calls that made its predictions better, which is the idea behind tool calling being trained into today's models.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Anthropic, Writing effective tools for agents: https://www.anthropic.com/engineering/writing-tools-for-agents
- Anthropic, Building effective agents (appendix on prompt-engineering tools): https://www.anthropic.com/engineering/building-effective-agents
- Claude tool use overview: https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
- Implementing tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
- Understanding JSON Schema: https://json-schema.org/understanding-json-schema
- Stripe, idempotent requests (the canonical explanation): https://docs.stripe.com/api/idempotent_requests
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.