rumblr Work in progressWIP

● The AI Primer · Lesson 41 · Part 2: building systems people rely on

Tools

design, validation and safety

You'll be able to explain Tool design, validation, idempotency, approvals, least privilege

Members · open during launch 21 min9 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. The model writes a request (tool_use). Your code validates it, runs it, and returns a tool_result.
  2. Validate shape (JSON Schema) and meaning (business rules). Return every problem, each with its fix.
  3. Descriptions are prompts: say what the tool does, when to use it, and when not to.
  4. Prefer fewer, higher-level tools: punishes long chains of calls.
  5. Past a few dozen tools, load only the relevant ones per request.
  6. Writes get idempotency keys; irreversible or high-value actions get human approval; credentials get the narrowest scopes.

Level 1

The practitioner's guide

In one sentence

A tool is a function you let the model ask for, and tool design is everything that decides whether the request it fills in is the right one, has correct arguments, and is safe to carry out: the definition the model reads, the checks your code runs, and the gates in front of the real system.

When you need it

The moment a model's output does something rather than says something: looks up a record, refunds a payment, sends an email. A chatbot that only answers from its weights has no tools and needs none of this. A one-off script with one read-only tool needs the definition and little else. Everything in this lesson becomes necessary as soon as a tool writes, costs money or can be called by a model that was misled. The tell: if a wrong tool call would be visible to a customer, an auditor or a bank, you need the checks. Two of this lesson's measurements show how much the design decides. The same three tools, described vaguely ("search stuff"), get the right tool for 2 of 6 requests; described precisely (what it does, when to use it, when not to), 6 of 6. And a job that takes five tool calls in a row, each 97% likely to be right, succeeds 85.9% of the time, so about one run in seven fails; as one high-level call it succeeds 97% of the time.

Your options

Six layers of protection between the model's request and the real system, from the cheapest to the most certain. They stack; a production registry has all of them.

Option What it does What it guarantees What it costs Where it lives
A schema in the definition A JSON Schema tells the model which fields exist, their types and which are required Nothing by itself; the model does its best to match A few tokens per tool per call Your tool definition
Strict mode The provider constrains the model's output, token by token, to your schema Arguments in exactly the schema's shape and a tool name that exists (Claude's strict tool use; OpenAI's strict mode) A schema written to the provider's rules (additionalProperties: false, required fields); the compiled grammar is cached The model server
Validation in your registry Checks shape and then business rules, and returns every problem with what a correct value looks like A wrong call never reaches the system, and the model can fix all its mistakes in one retry (four at once in the lesson's example) A validator and the rules; error messages worth writing Your code
Descriptions and consolidation Precise descriptions, few parameters, enums for closed sets, and one high-level tool in place of a chain The right tool chosen far more often (6 of 6 against 2 of 6 here), and fewer chances to fail in a row Writing, and an evaluation set to measure selection Your tool definitions
Loading only the relevant tools Retrieval over the catalogue picks the few tools a request needs (5 of 30 in the lesson), or the provider searches it on demand Selection accuracy holds as the catalogue grows; far fewer definition tokens per call One embedding or search per request; the top few most-used tools kept always loaded Your code, or the model server (Claude's tool search)
Safety gates Scopes per agent, a dry run that describes instead of acting, human approval over a threshold, an idempotency key on every write Least privilege, rehearsals, a person's sign-off, and a retry that cannot repeat a refund A credential model, an approver in the loop, a store of keys and results Your code, directly in front of the real system

How to choose

Decide by what the tool can do, not by how clever the model is.

  • Read-only lookups: a schema, a good description, and validation that returns actionable errors. Strict mode if the provider offers it; it removes a class of retries for free.
  • Anything that writes: all of the above, plus an idempotency key on the call and business-rule checks before it runs. C-999 fits the pattern ^C-\d+$ and is still a customer that does not exist.
  • Anything irreversible or over a money threshold: an approval rule, and a declined message that tells the model not to retry.
  • Any agent that reads untrusted text (email, web pages, documents): the narrowest scopes you can give it, so a tricked model cannot refund, not merely should not.
  • A catalogue past a few dozen tools: load per request, or route to a sub-agent that holds only its own tools. Claude's docs put the accuracy drop past 30 to 50 tools and recommend search from 10 tools up.
  • Whatever you pick, prefer fewer, higher-level tools. A chain of five calls compounds five chances to go wrong; the sequencing belongs in tested code, not in the model.

What it costs

Tokens: every definition is re-sent on every call, so a long catalogue is a standing charge. Claude's tool-search docs put a typical multi-server setup (GitHub, Slack, Sentry, Grafana, Splunk) at about 55,000 tokens of definitions before any work is done, and on-demand loading cuts that by over 85%. Tool results cost tokens too: Anthropic's tool-writing guide reports that a concise response format used about a third of the tokens of a detailed one, and that Claude Code caps a tool response at 25,000 tokens by default. Latency: validation is microseconds; an approval gate is however long a person takes, so the call parks as awaiting_approval rather than blocking. Quality: descriptions are the cheapest lever, and the lesson's 2-of-6 to 6-of-6 swing cost only words. Effort: the registry in this lesson, with every gate, is a few hundred lines, and the idempotency store can be a dictionary until it needs to survive a restart.

What breaks

  • Schema-valid, false. Strict mode and schemas prove shape, not truth. Keep the business-rule check, and return an error that says how to find the right id.
  • "Invalid input." A bare error makes the model guess. List every problem with the expected format ("date as YYYY-MM-DD"); the model fixes them all in one retry.
  • Vague descriptions. A model that cannot tell the tools apart picks the first one every time; in the lesson it got exactly the two requests that happened to belong to that tool.
  • Too many tools. Selection accuracy falls and definition tokens climb together. Load per request.
  • A retry that pays twice. A refund that times out after the bank processed it gets retried. Without a key the customer gets $40; with the same key the registry returns the stored result and refunds once. Stripe's API keeps a key's first result and returns it on every repeat.
  • A persistent model after a decline. "Declined" alone invites another attempt. Say "do not retry, tell the user".
  • One admin credential. If the agent is tricked, it can do anything the credential allows. Scope per tool and per agent.

In the wild

Claude's Messages API accepts strict: true on a tool and guarantees the arguments match the schema and the name is valid; its tool search tool defers tool definitions and loads them when the model searches for them; OpenAI's function calling has a strict mode its guide recommends always enabling. Anthropic's Writing effective tools for agents is the design reference behind sections 3 and 4 here: a few thoughtful tools over one per endpoint, consistent namespacing (asana_search, jira_search), responses that return meaning rather than identifiers, and evaluations built from dozens of real prompt and response pairs. Stripe's idempotent requests are the canonical form of the idempotency key: a client-chosen key, the first result saved and replayed, keys pruned after 24 hours. Toolformer (Schick et al., 2023) is the paper that showed a model can learn when and how to call an API, which is why today's models emit tool calls at all.

Go deeper

Level 2 writes the tool call out message by message, walks a call through every diamond of the registry in the order they are checked, measures the description experiment and the compounding formula, builds tool retrieval from cosine similarity, and replays the lost-reply refund with and without a key. If you only needed to choose, you are done.

Level 2

How it works, from scratch

A language model can only produce text. A tool is how that text turns into an action: looking up a record, refunding a payment, sending an email. This lesson builds a tool registry from scratch and shows the six things that decide whether tools work in production: how a call really happens, validation, descriptions, granularity, how many tools you expose, and safety.

Chapter 1

What a tool call really is

Everyday picture You hire a new assistant who is brilliant but may not touch anything. When they need something done, they fill in a request form ("refund customer C-100, $20") and hand it to you. You decide whether to carry it out, and you hand back a note saying what happened. The assistant is the model. The form is a tool call. You are your code.

Tiny worked example Here's the whole exchange for "what's 2 + 3?" with one tool, written as the actual messages:

you -> model   tools: [{"name": "add", "description": "Add two integers.",
                        "input_schema": {"type": "object",
                          "properties": {"a": {"type": "integer"}, "b": {"type": "integer"}},
                          "required": ["a", "b"]}}]
               user: "what's 2 + 3?"
model -> you   tool_use {"id": "toolu_1", "name": "add", "input": {"a": 2, "b": 3}}
               stop_reason: "tool_use"          <- "please run this for me"
you            (validate the input, run add(2, 3) -> 5)
you -> model   user: tool_result {"tool_use_id": "toolu_1", "content": "5"}
model -> you   "2 + 3 = 5."   stop_reason: "end_turn"

input_schema is a JSON Schema, a JSON document that describes the shape other JSON must have: which fields exist, what type each is, and which are required. It's the "form template" the model fills in.

Figure 1 · Diagram

Reading it: follow the arrows top to bottom. The model's only arrow towards the real system goes through your code. Nothing the model writes can touch the real system unless your code decides to act on it. That middle box, "check the form", is where the rest of this lesson lives.

The code ToolRegistry.register stores a Python function with its name, description and schema; ToolRegistry.definitions() produces the list you send to the model; ToolRegistry.call() runs a call through every check.

Why it matters The separation is the security model. A model can be wrong or tricked. Your code is where you enforce what's allowed.

Chapter 2

Validate every call before running it

Everyday picture A bank clerk checks a withdrawal slip twice. Is it filled in properly: every box, a real date, a positive amount? Then does it make sense: does this account exist, does it have the money? Only then do they open the drawer.

Tiny worked example A payment tool takes an amount, a currency (EUR, GBP or USD) and an optional pay_on date. The model sends {"ammount": 5, "currency": "usd", "pay_on": "next friday"}. The validator returns every problem at once, each one saying what a correct value looks like:

$: missing required field 'amount'
$: unexpected field 'ammount' (allowed: amount, currency, pay_on)
$.currency: must be one of ['EUR', 'GBP', 'USD'], got 'usd'
$.pay_on: 'next friday' does not match the required format. Expected: date as YYYY-MM-DD

Compare that to invalid input. With the first, the model fixes all four mistakes in one retry. With the second, it guesses.

Figure 2 · Diagram

Reading it: a call enters at the top and must pass every diamond to reach run the tool at the bottom. Each exit on the side is a named status in ToolOutcome, and each returns text the model can act on. The order is deliberate: permissions come before anything that reveals how the tool works, and cheap checks come before expensive ones. Approval and idempotency sit last, right next to the action they protect.

Why it matters "Structured output" and strict mode guarantee the shape of the arguments, not that they're true. "C-999" matches the pattern ^C-\d+$ perfectly and is still a customer that doesn't exist. Schema checks and business-rule checks are both required.

With the Claude API you can add "strict": true to a tool definition (definitions(strict=True) here). The API then guarantees the model's arguments validate against the schema. You still need the business-rule checks.

In code: validate checks a value against a JSON Schema subset and returns every problem with its path. ToolRegistry.call walks the diamonds above in order and wraps the result in a ToolOutcome, whose ToolOutcome.is_error says whether the model should treat it as a failure.

Chapter 3

Descriptions are prompts

Everyday picture A wall of drawers labelled "stuff", "things" and "items". Even a careful person opens the wrong one. Relabel them "Policies (PTO, travel, VPN). Not customers", and nobody hesitates.

Tiny worked example A stand-in model (pick_tool) chooses the tool whose name and description share the most words with the request. For "billing status for customer 1042", the vague descriptions share zero words with every tool (a three-way tie, so it guesses the first). The precise lookup_record description shares "billing", "status" and "customer" and wins.

Figure 3 · Drawn from the lesson's code

vague descriptions precise descriptions 0 1 2 3 4 5 6 requests routed to the right tool Same model, same requests, new descriptions 2/6 6/6

Vague descriptions pick the right tool for 2 of 6 requests; precise descriptions of the same tools get all 6 right

Reading it: two bars, one per set of descriptions, each out of the same six labelled requests. With vague descriptions the model gets 2 of 6, and only the two that happen to belong to the first tool, which it picks every time it can't tell them apart. With descriptions that say what each tool is for, when to use it and when not to, it gets 6 of 6. Only the descriptions changed.

Why it matters Real models are far better readers than word overlap, but they choose from exactly the same text. Precise descriptions, a few parameters, enums instead of free text, and error messages that explain the fix remove a large share of agent errors.

In code: selection_accuracy runs pick_tool over the labelled requests and counts the correct picks, which is what the figure plots for VAGUE_TOOLS and PRECISE_TOOLS.

Chapter 4

Fewer, higher-level tools

Everyday picture Asking an assistant to "book my trip to Denver" versus dictating five separate forms (find flight, hold seat, find hotel, reserve room, add to calendar). Every hand-off is another chance to drop something.

Figure 5 · Diagram

Reading it: on the left the model has to pick five tools in the right order and carry each output into the next input by hand. On the right the same work is one call, and the sequencing lives in ordinary tested code.

The chance that a chain of calls all succeed:

Level 3: the formula and its symbols

Symbols

Symbol Meaning
probability that one call is chosen and filled in correctly
number of calls in the chain

In words: multiply the per-call success rate by itself once per call.

On the example: with and , , so about 1 run in 7 fails. With one high-level call it's .

In Python:

p = 0.97
# five calls that must all succeed
round(p ** 5, 3)  # → 0.859
# about 1 run in 7 fails
round(1 - p ** 5, 2)  # → 0.14
# one high-level call
p ** 1  # → 0.97

Figure 4 · Drawn from the lesson's code

2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 calls in the chain (n) 0.0 0.2 0.4 0.6 0.8 1.0 P(every call succeeds) = p^n Reliability compounds: fewer, higher-level tools p = 0.90 per call p = 0.95 per call p = 0.97 per call p = 0.99 per call

Success falls with every extra call: at 99% per call twenty calls succeed about 82% of the time, at 90% ten calls barely a third

Reading it: the x-axis is how many calls the job takes, and the y-axis is the chance the whole job succeeds. Each curve is a different per-call reliability. Even at 99% per call, twenty calls succeed only about 82% of the time, and at 90% per call, ten calls succeed barely a third of the time. Fewer calls is the cheapest reliability you can buy.

In code: chain_success evaluates .

Chapter 5

Too many tools: load only the relevant ones

Everyday picture A warehouse versus a toolbox. You don't send a plumber into a warehouse of 500 tools; you hand them the five they need for today's job.

Tiny worked example A catalogue of 30 tools. For "I forgot my password and I'm locked out", select_tools embeds the request, compares it to every tool description, and sends the model only the top 5, with reset_password first.

Figure 7 · Diagram

Reading it: this is retrieval, the same machinery as RAG, but the "documents" are tool descriptions. The catalogue is embedded ahead of time; per request you embed one string and take the top matches.

Figure 6 · Drawn from the lesson's code

reset_password enroll_mfa check_vpn_status renew_device_certificate request_laptop report_phishing fix_printer submit_expense book_travel get_per_diem request_pto get_pto_balance report_sick_day request_parental_leave get_payslip get_salary_band start_onboarding list_invoices list_payments reconcile_invoices create_invoice send_email search_policies open_ticket close_ticket get_org_chart book_meeting_room order_office_supplies translate_text summarize_document I forgot my password and I'm locked out submit my car mileage expense how many vacation days are left suspicious email asking for my login reconcile vendor invoices for Q3 Cosine similarity: each request lights up a few tools −0.2 0.0 0.2 0.4 0.6 0.8

Each request is similar to only a small cluster of related tools and dim against the rest of the catalogue

Reading it: each row is a user request and each column a tool from the catalogue. Brighter cells mean more similar. Each request lights up a small cluster of related tools and stays dark everywhere else. Sending only that cluster keeps the model's choice small, which is why accuracy holds up as the catalogue grows.

Why it matters Selection accuracy drops as the tool list grows past a few dozen, and every definition costs input tokens on every call. The alternatives are routing the request to a sub-agent that holds only the relevant tools, or using a provider's built-in tool search.

Chapter 6

Safety: idempotency, dry runs, approval, least privilege

Idempotency. Everyday picture: pressing a lift button twice still takes you to the floor once. An operation is idempotent when doing it twice has the same effect as doing it once. Worked example: a refund call times out after the bank processed it, so the agent retries.

Figure 8 · Diagram

Reading it: the first reply is lost on the way back (the crossed arrow), so the agent can't know the refund happened and sensibly retries. Because the retry carries the same key, the registry returns the stored result instead of refunding again. Without the key, the customer gets $40.

Dry run. A rehearsal: run every check, then describe the action instead of doing it. Useful for previews and for shadow-mode rollouts (primer.agents.deployment).

Human approval. Everyday picture: a manager signs off on spending over a limit. Irreversible actions, or ones over a threshold such as refunds over $500, park as awaiting_approval until a person decides.

Figure 9 · Diagram

Reading it: the approval gate sits between the model's request and the action. Both branches return text the model can act on. "Do not retry" in the declined message matters, because otherwise a persistent model asks again.

Least privilege. Everyday picture: a valet key starts the car but won't open the boot. Give each agent credentials (scopes, named permissions such as payments:write) for its job only. An agent that reads tickets holds tickets:read, so even if a malicious email tricks it, it cannot issue refunds.

In code: a Tool carries its safety settings: the scopes it needs, whether it reads, writes or acts irreversibly, and an optional approval rule, which Tool.requires_approval combines. ToolRegistry.call enforces them, given the caller's credentials, an idempotency key, a dry-run flag and an approver.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: How do you design tools so the model picks the right one with correct arguments?Think it through, then reveal

A: Give each tool one clear job and a description that says what it does, when to use it and when not to. Keep parameters few, use enums for closed sets, and put the format in the schema description ("date as YYYY-MM-DD"). Prefer one high-level tool over several low-level ones. Validate every call and return actionable errors so the model can self-correct. Measure tool selection accuracy in evals, and load tools dynamically when there are many.

Question 2Q: The schema says customer_id must match ^C-\d+$. Is that enough validation?Think it through, then reveal

A: No. The schema checks shape, not truth. C-999 is well-formed and may not exist. Add a business-rule check before acting, and return an error that tells the model how to find the right id.

Question 3Q: A refund call timed out. Should the agent retry?Think it through, then reveal

A: Only if the call is idempotent. Send an idempotency key with every write, and the retry returns the original result instead of refunding twice.

Question 4Q: Why not give the agent one admin credential for everything?Think it through, then reveal

A: Blast radius. If the agent is tricked (for example by prompt injection in a document it reads), it can do anything the credential allows. Scope credentials per tool and per agent, and put irreversible actions behind human approval.

Primary sources

The papers behind this lesson

Schick et al., Toolformer: Language Models Can Teach Themselves to Use Tools (2023).

It showed a language model can learn when to call an API, which one and with what arguments, by keeping only the self-generated calls that made its predictions better, which is the idea behind tool calling being trained into today's models.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Writing effective tools for agents: https://www.anthropic.com/engineering/writing-tools-for-agents
  • Anthropic, Building effective agents (appendix on prompt-engineering tools): https://www.anthropic.com/engineering/building-effective-agents
  • Claude tool use overview: https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
  • Implementing tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
  • Understanding JSON Schema: https://json-schema.org/understanding-json-schema
  • Stripe, idempotent requests (the canonical explanation): https://docs.stripe.com/api/idempotent_requests

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.