The lesson in one minute
What you'll be able to explain
- A model call sends the full list of messages and returns one message; the model keeps no state between calls.
- A tool call is a structured request (
tool_use). Your code validates it, runs it, and returns atool_resultwith the matching id. The model executes nothing. stop_reasontells you what to do next:tool_usemeans run tools and call again,end_turnmeans done.- Every call re-sends the whole history, so input cost grows quadratically with the number of steps in a loop.
Level 1
The practitioner's guide
In one sentence
Talking to a model means sending it the whole
conversation so far as a list of messages and getting one message back, and
tool calling is the case where that message is a structured request ("run
get_weather with {"city": "Paris"}") that your code may carry out and
answer.
When you need it
Every system built on a language model does this,
from a one-shot classifier to a multi-agent research team, so the message
format is not optional. Tool calling is. You need it the moment the model
must reach past its own weights: look something up, run a query, compute a
number, send an email, change a record. You don't need it when the answer is
words the model already knows, and you don't need it when one structured
answer is enough (a label, a JSON row): that is structured output
(primer.ml.structured_output), which is a tool call without the "run it and
come back" step. The tell: if your code reads the model's prose to decide
which function to call, or scrapes a city name out of a sentence, you need
tool calling. The model was trained to hand you that request as data.
Two facts about the exchange decide most of what follows, and both come from
this lesson's traced example. First, the model executes nothing: the reply
to "What's the weather in Paris?" is a tool_use block asking for
get_weather, and the 18°C comes from your code running the function and
sending the result back. Second, the model keeps no state between calls, so
every call carries the whole history again. In the lesson's ten-call loop,
that turns a conversation 7,000 tokens long into 42,500 input tokens billed.
Your options
Five ways to get an action out of a model, from the cheapest to the most certain:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| A scripted stand-in | A function you write plays the model, replying by rule | The exact reply you scripted, every run, offline | Nothing per call; proves nothing about a real model | Your tests and demos |
| A local open model, text only | Weights on your own machine answer in plain text | Privacy: nothing leaves the box; no per-token bill | Hardware, smaller models, and your own parsing if you need actions | Your machine (Ollama and similar) |
| Text you parse | Ask the model to write the action in prose, then match it with a regular expression | Nothing; it works until the wording drifts | A parser you maintain and a retry loop | Your code |
| Native tool calling | The model replies with a typed tool_use block: a name, an id, and the arguments as a parsed JSON object |
A request in your schema's shape; with strict mode, exactly your schema | Tokens for the tool definitions, a tool-use system prompt the API adds (286 tokens on Claude Opus 5, per the pricing page), and one extra round trip per call | The model server produces it; your code runs it |
| Server-side tools | The provider runs the tool (web search, web fetch, code execution) and returns the result in the same reply | The result arrives with no handler code on your side | Per-use fees on some tools (Claude's web search is $10 per 1,000 searches) and no control over execution | The model server |
How to choose
Start from what has to happen, and who is allowed to make it happen.
- Testing the code around the model (a loop, a budget, an approval step): the scripted stand-in. It reproduces a model that loops, calls the wrong tool or follows an injected instruction on cue, which a real model rarely does when you want it to.
- Private data, no budget, or a laptop on a plane: a local open model. Keep it to text unless the model and server both support tool calling (Ollama does, for models trained for it; this lesson's local adapter is text only).
- Anything a program acts on: native tool calling, never parsed prose. Turn on strict mode when the arguments feed straight into code.
- Web search or sandboxed code that you would otherwise host yourself: a server-side tool, if its fees and its lack of a hook for your own checks suit you.
- Whatever you pick, your code stays the only thing that acts. Validate the arguments, check permissions, then run the tool; the model only asks.
What it costs
Input tokens are billed on the whole history, every call,
so the cost of a loop grows with the square of its length: the lesson's ten
calls bill 42,500 tokens, six times the 7,000-token conversation they end
with, and the tenth call alone re-sends 6,500. At Claude Opus 5's list price
of $5 per million input tokens (the pricing page, fetched for this guide)
that is about 21 cents per ten-step task before any output; on Haiku 4.5 at
$1 per million, about 4 cents. Prompt caching sells re-read tokens at a
tenth of the price, which is why it is the first lever on any loop
(primer.agents.cost). Tool definitions ride along on every call too, so a
long tool list is a standing charge. Latency is one network round trip per
model call plus the tool's own time, and a tool call always means at least
two model calls.
What breaks
- Dropping the assistant turn. Append the exact content blocks the model returned, not just their text: the ids in them are how the next call matches results to requests, and real APIs can include blocks (thinking) that must go back unchanged.
- Unmatched ids. A
tool_resultwhosetool_use_idmatches no request answers nothing. Return one result per request, and all of them in a single user message when the model asked for several. - Swallowing failures. A tool that throws and returns nothing leaves the model waiting. Send a result flagged as an error with a message the model can act on, so it can retry or ask.
- Trusting the request. Arguments arrive as a parsed object in the shape of your schema, not as safe values. Validate them and check permissions before running anything; the model is never a security boundary.
- Ignoring
stop_reason.tool_usemeans run tools and call again;end_turnmeans done;max_tokensmeans the reply was cut off and its JSON may be incomplete;refusalmeans the model declined. A loop that checks only for text gets all four wrong. - Missing arguments guessed. Asked for the weather with no city, a model may invent one rather than ask (Claude's tool-use docs call the asking behaviour model-dependent and not guaranteed). Make optional what is optional, and validate the rest.
In the wild
Claude's Messages API is the shape this lesson uses
directly: tool_use and tool_result blocks, stop_reason, parallel tool
calls, strict: true for schema-exact arguments, and server tools (web
search, web fetch, code execution) alongside the client tools you run. Its
SDKs add a tool runner that drives the request, run, reply loop for you.
OpenAI's function calling is the same five-step round trip in different
field names, with its own strict mode; the guide recommends turning it on.
Ollama exposes tool calling for local models, where again the model returns
the call and your code runs it. Every agent framework, from the smallest
loop to a multi-agent system, is built on this exchange, and the papers
behind it (Toolformer, ReAct) are in this lesson's papers section.
Go deeper
Level 2 traces the four messages of one tool call by hand, lays out the message format field by field, shows the one interface that a scripted model and a real one both implement, and derives the quadratic cost of a loop with a formula you can rerun. If you only needed to choose, you are done.
Level 2
How it works, from scratch
Every applied-AI system in this primer, from a one-shot classifier to a multi-agent research team, talks to the model the same way: it sends a list of messages and gets one message back. This lesson shows that exchange exactly, and then the one feature that turns a chatbot into an agent: tool calling.
Chapter 1
The everyday picture
Think of the model as a brilliant consultant who works by post. You mail a letter with your whole conversation so far (the consultant keeps no notes between letters) and get one letter back. The reply is either an answer, or a filled-in request form: "please look up the weather in Paris and send me the result." The consultant never picks up the phone themselves. You decide whether to run the request, run it, and mail back the result, together with the whole conversation again.
Two facts fall straight out of this picture, and they explain most of the behaviour of real systems:
- The model executes nothing. A tool call is a request. Your code is the only thing that acts, so your code is where safety lives.
- Every call resends everything. The model has no memory between calls, so each request carries the full history. Long conversations cost more on every single turn.
Chapter 2
A tiny worked example: one tool call, traced
A user asks "What's the weather in Paris?" and the application offers one
tool, get_weather. Four messages later, the user has an answer:
| # | Role | Content | Who produced it |
|---|---|---|---|
| 1 | user | "What's the weather in Paris?" | the user |
| 2 | assistant | tool_use id=toolu_1, name=get_weather, input={"city": "Paris"} |
the model (1st call) |
| 3 | user | tool_result for toolu_1: "18°C, sunny" |
your code, after running the tool |
| 4 | assistant | "It's 18°C and sunny in Paris." | the model (2nd call) |
Message 3 has role user even though no human typed it: tool results always
travel back in the user's turn. The id in message 2 and the tool_use_id
in message 3 match, which is how the model knows which request a result
answers when it asked for several at once. trace_tool_call_round_trip()
produces exactly this transcript.
Figure 1 · Diagram
sequenceDiagram
participant U as User
participant A as Your code
participant M as Model
participant T as get_weather tool
U->>A: What's the weather in Paris?
A->>M: messages [1] + tool definitions
M-->>A: tool_use get_weather {city: Paris} (stop_reason: tool_use)
A->>A: validate the arguments, check permissions
A->>T: get_weather("Paris")
T-->>A: 18°C, sunny
A->>M: messages [1, 2, 3 = tool_result]
M-->>A: It's 18°C and sunny in Paris. (stop_reason: end_turn)
A-->>U: It's 18°C and sunny in Paris.
get_weather starts at your code, and the self-arrow on your
code ("validate, check permissions") is where production systems put their
guardrails. Notice too that the second call to the model carries messages
1 to 3, not just the new result: the model is stateless, so the history
travels every time.In code: ToolCall holds the request in message 2, and
tool_result_block builds the reply in message 3, carrying the matching id.
Chapter 3
The message format
Messages use the Anthropic Messages API shape directly, so what you learn here maps one-to-one onto production code:
{"role": "user", "content": "What's our PTO policy?"}
{"role": "assistant", "content": [
{"type": "text", "text": "Let me look that up."},
{"type": "tool_use", "id": "toolu_1", "name": "search_kb", "input": {"query": "PTO"}},
]}
{"role": "user", "content": [
{"type": "tool_result", "tool_use_id": "toolu_1", "content": "hr-001: ..."},
]}
| Field | Meaning |
|---|---|
role |
user (the human, or your code returning tool results) or assistant (the model) |
content |
a plain string, or a list of content blocks |
text block |
ordinary words |
tool_use block |
the model asking for a tool: id, name, and input (a JSON object matching the tool's schema) |
tool_result block |
your answer to one tool_use, matched by tool_use_id; set is_error: true when the tool failed |
stop_reason |
why the reply ended: end_turn (finished), tool_use (wants a tool), max_tokens (ran out of room), refusal (declined) |
A tool definition is a name, a description and a JSON Schema for the
input. The model chooses tools by reading those descriptions, so writing
them well is prompt engineering (see primer.agents.tools).
In code: LLMResponse is one reply, normalized: its text, its
ToolCall list, its stop reason, its Usage and the exact content blocks
to append as the assistant turn. content_blocks, last_user_text,
tool_results and tool_calls_so_far read a conversation in this format.
Chapter 4
Two implementations of one interface
Every agent lesson is written against one small interface, LLM, with one
method, LLM.complete(system=..., messages=..., tools=...). Two classes
implement it:
Figure 2 · Diagram
flowchart LR L["LLM interface<br/>complete(system, messages, tools)"] L --> S["ScriptedLLM<br/>a Python function decides the reply<br/>offline, instant, deterministic"] L --> C["ClaudeLLM<br/>the real model via the anthropic SDK<br/>needs an API key"] S --> D["demos and tests<br/>reproduce any failure on purpose"] C --> P["production<br/>same agent code, real behaviour"]
ScriptedLLM lets every lesson run offline and
deterministically, and lets tests stage a model that loops, calls the wrong
tool or follows an injected instruction, which is hard to get a real model
to do on cue. Swapping in ClaudeLLM runs the identical loop against the
real thing.In code: ClaudeLLM.complete builds its request with claude_request
and turns the API's reply into an LLMResponse. OllamaLLM is a third
implementation for a local open model (text only), using ollama_request
and parse_ollama_reply.
Chapter 5
Cost: why every call pays for the whole conversation
Because the model is stateless, input tokens are billed per call on the entire history. In an agent loop with calls, where each call adds about new tokens to a history that started at tokens, the total input billed is:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| number of model calls in the loop | 10 | |
| tokens in the first request (system prompt, tools, question) | 2,000 | |
| tokens each step adds (the model's tool call plus the tool's result) | 500 | |
| which call we're on, 1 to | ||
| add up the cost of every call | ||
| 0 + 1 + … + (n − 1): how many "steps of history" pile up in total | 45 |
In words: "each call pays for the starting prompt plus everything added so far, so the history term grows with the square of the number of steps."
With the numbers: 10 × 2,000 + 500 × 45 = 20,000 + 22,500 = 42,500 input tokens for a task whose final conversation is only 7,000 tokens long: the tenth call re-sends 6,500 tokens, and its own 500-token step brings the history to 7,000.
Level 3: in Python
n, h_0, t = 10, 2000, 500
# call i re-sends h_0 and i - 1 steps
calls = [h_0 + (i - 1) * t for i in range(1, n + 1)]
calls[0], calls[-1] # → (2000, 6500)
# Σ over every call
sum(calls) # → 42500
# the shortcut on the right agrees
n * h_0 + t * n * (n - 1) // 2 # → 42500
Figure 3 · Drawn from the lesson's code
Ten calls bill 42,500 input tokens in total, six times the 7,000-token conversation they end with, because every call re-sends the history
primer.agents.cost, primer.agents.context).In code: loop_input_tokens lists the tokens billed on each call of
such a loop. estimate_tokens is the four-characters-per-token rule of
thumb, and conversation_chars measures everything a call resends, which is
how ScriptedLLM gives its fake replies realistic Usage.
Test yourself
5 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1What exactly happens when a model "uses a tool"?Think it through, then reveal
You send tool definitions (name, description, JSON Schema) with the
messages. The model replies with a tool_use block and stop_reason: tool_use. Your code validates the arguments, runs the function, and sends a
new request with the history plus a tool_result block carrying the same
id. The model then answers or asks for another tool.
Question 2Why is the model never a security boundary for tool use?Think it through, then reveal
It only produces requests. Everything that actually happens goes through your code, so validation, permissions, approvals and rate limits all belong there, and a manipulated model can only ever ask.
Question 3The model asks for three tools in one reply. How do you send the results?Think it through, then reveal
Run them (concurrently if independent) and return all three tool_result
blocks in a single user message, each with its own tool_use_id. For one
that failed, return a tool_result with is_error: true and an actionable
message rather than dropping it.
Question 4A 20-step agent loop costs far more than 20 times a single call. Why?Think it through, then reveal
Each call re-sends the full, growing history, so the input billed is a sum that grows with the square of the number of steps. Caching the stable prefix and trimming tool outputs attack exactly this.
Question 5Why test agents against a scripted model at all?Think it through, then reveal
Real models are non-deterministic and rarely misbehave on cue. A scripted model reproduces loops, bad arguments and injected instructions exactly, so the code that must handle them can be tested every time.
Primary sources
The papers behind this lesson
Showed a model can learn when to call external tools and how to use their results.
Read the annotated companion →The paper ↗The pattern of interleaving reasoning with tool calls that modern agent loops descend from.
Read the annotated companion →The paper ↗Researcher's shelf
Further reading
- Tool use overview: https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
- Implementing tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
- Messages API reference: https://docs.claude.com/en/api/messages
- Python SDK: https://github.com/anthropics/anthropic-sdk-python
- Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.