rumblr Work in progressWIP

● The AI Primer · Lesson 38 · Part 2: building systems people rely on

Talking to a model

messages, and what tool calling really is

You'll be able to explain The message format, and what tool calling really is

Members · open during launch 16 min3 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. A model call sends the full list of messages and returns one message; the model keeps no state between calls.
  2. A tool call is a structured request (tool_use). Your code validates it, runs it, and returns a tool_result with the matching id. The model executes nothing.
  3. stop_reason tells you what to do next: tool_use means run tools and call again, end_turn means done.
  4. Every call re-sends the whole history, so input cost grows quadratically with the number of steps in a loop.

Level 1

The practitioner's guide

In one sentence

Talking to a model means sending it the whole conversation so far as a list of messages and getting one message back, and tool calling is the case where that message is a structured request ("run get_weather with {"city": "Paris"}") that your code may carry out and answer.

When you need it

Every system built on a language model does this, from a one-shot classifier to a multi-agent research team, so the message format is not optional. Tool calling is. You need it the moment the model must reach past its own weights: look something up, run a query, compute a number, send an email, change a record. You don't need it when the answer is words the model already knows, and you don't need it when one structured answer is enough (a label, a JSON row): that is structured output (primer.ml.structured_output), which is a tool call without the "run it and come back" step. The tell: if your code reads the model's prose to decide which function to call, or scrapes a city name out of a sentence, you need tool calling. The model was trained to hand you that request as data.

Two facts about the exchange decide most of what follows, and both come from this lesson's traced example. First, the model executes nothing: the reply to "What's the weather in Paris?" is a tool_use block asking for get_weather, and the 18°C comes from your code running the function and sending the result back. Second, the model keeps no state between calls, so every call carries the whole history again. In the lesson's ten-call loop, that turns a conversation 7,000 tokens long into 42,500 input tokens billed.

Your options

Five ways to get an action out of a model, from the cheapest to the most certain:

Option What it does What it guarantees What it costs Where it lives
A scripted stand-in A function you write plays the model, replying by rule The exact reply you scripted, every run, offline Nothing per call; proves nothing about a real model Your tests and demos
A local open model, text only Weights on your own machine answer in plain text Privacy: nothing leaves the box; no per-token bill Hardware, smaller models, and your own parsing if you need actions Your machine (Ollama and similar)
Text you parse Ask the model to write the action in prose, then match it with a regular expression Nothing; it works until the wording drifts A parser you maintain and a retry loop Your code
Native tool calling The model replies with a typed tool_use block: a name, an id, and the arguments as a parsed JSON object A request in your schema's shape; with strict mode, exactly your schema Tokens for the tool definitions, a tool-use system prompt the API adds (286 tokens on Claude Opus 5, per the pricing page), and one extra round trip per call The model server produces it; your code runs it
Server-side tools The provider runs the tool (web search, web fetch, code execution) and returns the result in the same reply The result arrives with no handler code on your side Per-use fees on some tools (Claude's web search is $10 per 1,000 searches) and no control over execution The model server

How to choose

Start from what has to happen, and who is allowed to make it happen.

  • Testing the code around the model (a loop, a budget, an approval step): the scripted stand-in. It reproduces a model that loops, calls the wrong tool or follows an injected instruction on cue, which a real model rarely does when you want it to.
  • Private data, no budget, or a laptop on a plane: a local open model. Keep it to text unless the model and server both support tool calling (Ollama does, for models trained for it; this lesson's local adapter is text only).
  • Anything a program acts on: native tool calling, never parsed prose. Turn on strict mode when the arguments feed straight into code.
  • Web search or sandboxed code that you would otherwise host yourself: a server-side tool, if its fees and its lack of a hook for your own checks suit you.
  • Whatever you pick, your code stays the only thing that acts. Validate the arguments, check permissions, then run the tool; the model only asks.

What it costs

Input tokens are billed on the whole history, every call, so the cost of a loop grows with the square of its length: the lesson's ten calls bill 42,500 tokens, six times the 7,000-token conversation they end with, and the tenth call alone re-sends 6,500. At Claude Opus 5's list price of $5 per million input tokens (the pricing page, fetched for this guide) that is about 21 cents per ten-step task before any output; on Haiku 4.5 at $1 per million, about 4 cents. Prompt caching sells re-read tokens at a tenth of the price, which is why it is the first lever on any loop (primer.agents.cost). Tool definitions ride along on every call too, so a long tool list is a standing charge. Latency is one network round trip per model call plus the tool's own time, and a tool call always means at least two model calls.

What breaks

  • Dropping the assistant turn. Append the exact content blocks the model returned, not just their text: the ids in them are how the next call matches results to requests, and real APIs can include blocks (thinking) that must go back unchanged.
  • Unmatched ids. A tool_result whose tool_use_id matches no request answers nothing. Return one result per request, and all of them in a single user message when the model asked for several.
  • Swallowing failures. A tool that throws and returns nothing leaves the model waiting. Send a result flagged as an error with a message the model can act on, so it can retry or ask.
  • Trusting the request. Arguments arrive as a parsed object in the shape of your schema, not as safe values. Validate them and check permissions before running anything; the model is never a security boundary.
  • Ignoring stop_reason. tool_use means run tools and call again; end_turn means done; max_tokens means the reply was cut off and its JSON may be incomplete; refusal means the model declined. A loop that checks only for text gets all four wrong.
  • Missing arguments guessed. Asked for the weather with no city, a model may invent one rather than ask (Claude's tool-use docs call the asking behaviour model-dependent and not guaranteed). Make optional what is optional, and validate the rest.

In the wild

Claude's Messages API is the shape this lesson uses directly: tool_use and tool_result blocks, stop_reason, parallel tool calls, strict: true for schema-exact arguments, and server tools (web search, web fetch, code execution) alongside the client tools you run. Its SDKs add a tool runner that drives the request, run, reply loop for you. OpenAI's function calling is the same five-step round trip in different field names, with its own strict mode; the guide recommends turning it on. Ollama exposes tool calling for local models, where again the model returns the call and your code runs it. Every agent framework, from the smallest loop to a multi-agent system, is built on this exchange, and the papers behind it (Toolformer, ReAct) are in this lesson's papers section.

Go deeper

Level 2 traces the four messages of one tool call by hand, lays out the message format field by field, shows the one interface that a scripted model and a real one both implement, and derives the quadratic cost of a loop with a formula you can rerun. If you only needed to choose, you are done.

Level 2

How it works, from scratch

Every applied-AI system in this primer, from a one-shot classifier to a multi-agent research team, talks to the model the same way: it sends a list of messages and gets one message back. This lesson shows that exchange exactly, and then the one feature that turns a chatbot into an agent: tool calling.

Chapter 1

The everyday picture

Think of the model as a brilliant consultant who works by post. You mail a letter with your whole conversation so far (the consultant keeps no notes between letters) and get one letter back. The reply is either an answer, or a filled-in request form: "please look up the weather in Paris and send me the result." The consultant never picks up the phone themselves. You decide whether to run the request, run it, and mail back the result, together with the whole conversation again.

Two facts fall straight out of this picture, and they explain most of the behaviour of real systems:

  1. The model executes nothing. A tool call is a request. Your code is the only thing that acts, so your code is where safety lives.
  2. Every call resends everything. The model has no memory between calls, so each request carries the full history. Long conversations cost more on every single turn.

Chapter 2

A tiny worked example: one tool call, traced

A user asks "What's the weather in Paris?" and the application offers one tool, get_weather. Four messages later, the user has an answer:

# Role Content Who produced it
1 user "What's the weather in Paris?" the user
2 assistant tool_use id=toolu_1, name=get_weather, input={"city": "Paris"} the model (1st call)
3 user tool_result for toolu_1: "18°C, sunny" your code, after running the tool
4 assistant "It's 18°C and sunny in Paris." the model (2nd call)

Message 3 has role user even though no human typed it: tool results always travel back in the user's turn. The id in message 2 and the tool_use_id in message 3 match, which is how the model knows which request a result answers when it asked for several at once. trace_tool_call_round_trip() produces exactly this transcript.

Figure 1 · Diagram

Reading it: time runs downwards. Solid arrows are requests and dashed arrows are replies. Notice that the model never talks to the tool: every arrow into get_weather starts at your code, and the self-arrow on your code ("validate, check permissions") is where production systems put their guardrails. Notice too that the second call to the model carries messages 1 to 3, not just the new result: the model is stateless, so the history travels every time.

In code: ToolCall holds the request in message 2, and tool_result_block builds the reply in message 3, carrying the matching id.

Chapter 3

The message format

Messages use the Anthropic Messages API shape directly, so what you learn here maps one-to-one onto production code:

{"role": "user", "content": "What's our PTO policy?"}
{"role": "assistant", "content": [
    {"type": "text", "text": "Let me look that up."},
    {"type": "tool_use", "id": "toolu_1", "name": "search_kb", "input": {"query": "PTO"}},
]}
{"role": "user", "content": [
    {"type": "tool_result", "tool_use_id": "toolu_1", "content": "hr-001: ..."},
]}
Field Meaning
role user (the human, or your code returning tool results) or assistant (the model)
content a plain string, or a list of content blocks
text block ordinary words
tool_use block the model asking for a tool: id, name, and input (a JSON object matching the tool's schema)
tool_result block your answer to one tool_use, matched by tool_use_id; set is_error: true when the tool failed
stop_reason why the reply ended: end_turn (finished), tool_use (wants a tool), max_tokens (ran out of room), refusal (declined)

A tool definition is a name, a description and a JSON Schema for the input. The model chooses tools by reading those descriptions, so writing them well is prompt engineering (see primer.agents.tools).

In code: LLMResponse is one reply, normalized: its text, its ToolCall list, its stop reason, its Usage and the exact content blocks to append as the assistant turn. content_blocks, last_user_text, tool_results and tool_calls_so_far read a conversation in this format.

Chapter 4

Two implementations of one interface

Every agent lesson is written against one small interface, LLM, with one method, LLM.complete(system=..., messages=..., tools=...). Two classes implement it:

Figure 2 · Diagram

Reading it: the agent code on the right-hand side never knows which box it's talking to. ScriptedLLM lets every lesson run offline and deterministically, and lets tests stage a model that loops, calls the wrong tool or follows an injected instruction, which is hard to get a real model to do on cue. Swapping in ClaudeLLM runs the identical loop against the real thing.

In code: ClaudeLLM.complete builds its request with claude_request and turns the API's reply into an LLMResponse. OllamaLLM is a third implementation for a local open model (text only), using ollama_request and parse_ollama_reply.

Chapter 5

Cost: why every call pays for the whole conversation

Because the model is stateless, input tokens are billed per call on the entire history. In an agent loop with calls, where each call adds about new tokens to a history that started at tokens, the total input billed is:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
number of model calls in the loop 10
tokens in the first request (system prompt, tools, question) 2,000
tokens each step adds (the model's tool call plus the tool's result) 500
which call we're on, 1 to
add up the cost of every call
0 + 1 + … + (n − 1): how many "steps of history" pile up in total 45

In words: "each call pays for the starting prompt plus everything added so far, so the history term grows with the square of the number of steps."

With the numbers: 10 × 2,000 + 500 × 45 = 20,000 + 22,500 = 42,500 input tokens for a task whose final conversation is only 7,000 tokens long: the tenth call re-sends 6,500 tokens, and its own 500-token step brings the history to 7,000.

Level 3: in Python
n, h_0, t = 10, 2000, 500
# call i re-sends h_0 and i - 1 steps
calls = [h_0 + (i - 1) * t for i in range(1, n + 1)]
calls[0], calls[-1]  # → (2000, 6500)
# Σ over every call
sum(calls)  # → 42500
# the shortcut on the right agrees
n * h_0 + t * n * (n - 1) // 2  # → 42500

Figure 3 · Drawn from the lesson's code

2 4 6 8 10 model call in the loop 0 5000 10000 15000 20000 25000 30000 35000 40000 input tokens Every call re-sends the whole history running total final conversation size (billed once?) input tokens billed on this call

Ten calls bill 42,500 input tokens in total, six times the 7,000-token conversation they end with, because every call re-sends the history

Reading it: the bars are the input tokens billed on each call. They grow by the same amount every step, because each call re-sends the whole history. The line is the running total, and it curves upwards (quadratic growth). The dashed line is what you might naively expect: the size of the final conversation, billed once. The gap between the dashed line and the curve is why prompt caching, trimming tool outputs and keeping loops short are the main cost levers (primer.agents.cost, primer.agents.context).

In code: loop_input_tokens lists the tokens billed on each call of such a loop. estimate_tokens is the four-characters-per-token rule of thumb, and conversation_chars measures everything a call resends, which is how ScriptedLLM gives its fake replies realistic Usage.

Test yourself

5 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1What exactly happens when a model "uses a tool"?Think it through, then reveal

You send tool definitions (name, description, JSON Schema) with the messages. The model replies with a tool_use block and stop_reason: tool_use. Your code validates the arguments, runs the function, and sends a new request with the history plus a tool_result block carrying the same id. The model then answers or asks for another tool.

Question 2Why is the model never a security boundary for tool use?Think it through, then reveal

It only produces requests. Everything that actually happens goes through your code, so validation, permissions, approvals and rate limits all belong there, and a manipulated model can only ever ask.

Question 3The model asks for three tools in one reply. How do you send the results?Think it through, then reveal

Run them (concurrently if independent) and return all three tool_result blocks in a single user message, each with its own tool_use_id. For one that failed, return a tool_result with is_error: true and an actionable message rather than dropping it.

Question 4A 20-step agent loop costs far more than 20 times a single call. Why?Think it through, then reveal

Each call re-sends the full, growing history, so the input billed is a sum that grows with the square of the number of steps. Caching the stable prefix and trimming tool outputs attack exactly this.

Question 5Why test agents against a scripted model at all?Think it through, then reveal

Real models are non-deterministic and rarely misbehave on cue. A scripted model reproduces loops, bad arguments and injected instructions exactly, so the code that must handle them can be tested every time.

Primary sources

The papers behind this lesson

Schick et al., Toolformer: Language Models Can Teach Themselves to Use Tools (2023)

Showed a model can learn when to call external tools and how to use their results.

Read the annotated companion →The paper ↗
Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models (2022)

The pattern of interleaving reasoning with tool calls that modern agent loops descend from.

Read the annotated companion →The paper ↗

Researcher's shelf

Further reading

  • Tool use overview: https://docs.claude.com/en/docs/agents-and-tools/tool-use/overview
  • Implementing tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
  • Messages API reference: https://docs.claude.com/en/api/messages
  • Python SDK: https://github.com/anthropics/anthropic-sdk-python
  • Anthropic, Building effective agents: https://www.anthropic.com/engineering/building-effective-agents

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.