The lesson in one minute
What you'll be able to explain
- A trace records one run as a tree of spans: the task, then each model call, tool call and retrieval, with inputs, outputs, tokens, latency, cost, and model and prompt versions.
- Use the OpenTelemetry GenAI attribute names so traces flow into existing monitoring or LLM tracing tools without custom glue.
- Debug from traces: find the first error span; the fix often belongs in a tool or in retrieval, not the prompt.
- Link traces to feedback and evals: production failure → trace → golden case → fix → CI → deploy.
Level 1
The practitioner's guide
In one sentence
Observability for an agent means recording every run as a tree of timed steps (the task, then each model call, tool call and retrieval, with what went in and what came out), so that "why did it do that?" is a lookup and not a guess.
When you need it
From the first day real people use the system. The
tell: a user reports a wrong answer from yesterday and the only evidence is
the answer itself. This lesson's worked example is exactly that case. A user
asks "Where is order A-100?" (with a hyphen) and gets "I couldn't find order
A-100." From the answer alone the order might not exist; the trace shows a
tool error on one line, unknown order id A-100, and tells you the fix
belongs in the tool that doesn't normalise hyphens, not in the prompt. The
second tell is a bill you can't explain: in the lesson's six-order run,
input tokens climb with every model call because each call resends the
growing history, and only a per-call record makes that visible before the
invoice does. You don't need a tracing platform for a script you run by
hand, but even there, one structured record per model call (model, tokens,
latency, cost) pays for itself the first time something is slow.
Your options
Six ways to see inside a run, from the cheapest to the most complete:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Print statements and plain logs | Lines of text as the agent runs | You can read them, once; nothing ties one run's steps together | Nothing up front; hours later, when you need to find one run among thousands | Your code |
| Structured per-call records | One record per model call with model, tokens, latency, cost, prompt version | Dashboards for cost and volume; the bill becomes explainable | A logging library and a few fields per call | Your code |
| Traces with spans | Each step is a timed span with attributes and a parent; a run is a tree under one trace id | Replay any run; find the first error; see where the time goes | Instrumenting each call (a with block), and somewhere to store the trees |
Your agent loop |
| A vendor's own SDK | The tracing platform's client records the spans for you | Fastest start | Lock-in: changing vendors is a code change | Your code, tied to one service |
| OpenTelemetry with the GenAI conventions | The same spans in standard names (gen_ai.request.model, gen_ai.usage.input_tokens, gen_ai.tool.name) and the standard OTLP format |
Flows into the monitoring the organisation already runs and into any LLM tool; changing vendors is a collector setting | Learning the conventions; running a collector | Exporter and collector |
| An LLM tracing platform on top | Traces, user feedback, prompt versions, datasets and evals in one place | The loop from a thumbs-down to a golden test is a few clicks | Hosting it yourself or sending traces (with user data in them) to a service | A service |
How to choose
Decide by what you'll need to answer, then by who else needs to read it.
- You need to explain a bill: structured per-call records with tokens, model and cost; a sum over them is the invoice.
- You need to explain a decision: traces. The first error span, or the first surprising output, shows where the chain broke. Nothing less lets you replay a run.
- The organisation already runs Datadog, Grafana or Honeycomb: emit
OpenTelemetry spans with the GenAI attribute names, and let the collector
fan them out. Name your own attributes with a prefix (
app.prompt.version) so they never collide with the standard's. - You want traces linked to feedback and evals: put an LLM tracing platform behind the collector; most of them accept OpenTelemetry, so the agent code doesn't change.
- Whatever you pick, redact personal data before export, restrict who can read traces, set a retention limit, and keep tenants' traces apart. A trace holds the user's words.
What it costs
Instrumentation is cheap in code: a span is a with
block around a call, and an exception inside it marks the span as an error
by itself. It is not free in storage: a trace keeps every prompt and tool
result, so this lesson's three-step run (two model calls, one tool call,
900 ms on its toy clock, 195 input and 38 output tokens, about a tenth of
a cent) is a few kilobytes, and a million runs a day is gigabytes.
Distributed tracing at scale has always answered that by sampling: Google's
Dapper, the design OpenTelemetry descends from, kept overhead low by
tracing a sample of requests through shared libraries rather than every
request. Keep every trace while volume is small; sample when it isn't, but
always keep the ones a user flagged or a monitor caught. Latency cost is
negligible when export is asynchronous. The cost that matters is what you
learn from the trace itself: in the lesson's six-order run, model calls
take far more of the time than tool calls, so the latency wins are fewer
model calls and smaller contexts, not faster tools.
What breaks
- A log with no thread through it. Thousands of lines from thousands of runs, and no way to gather one run back together. A trace id on every span is the fix.
- Fixing the prompt when the tool was wrong. Without the trace, the hyphen bug above looks like a model mistake. Read the first error span before editing anything.
- Personal data in the trace store. Prompts and outputs hold names, emails and card numbers. Redact before export, and limit retention.
- Vendor-shaped traces. Attributes named after one product's SDK don't flow anywhere else. Use the standard names and a collector.
- Gaps between spans. Time outside any span (a queue, a retry wait, untraced code) is invisible in the tree. Instrument the retries and the guardrail checks too.
- Traces nobody reads. Recording is half the job. Store user feedback against the trace id and route flagged traces to a person, so each one can become a golden case.
In the wild
OpenTelemetry is the open standard for traces, metrics and
logs; its concepts page defines a trace as the path of a request through an
application and a span as a unit of work with a name, a parent, start and
end times, attributes, events and a status, which is exactly what this
lesson's Span holds. Its GenAI semantic conventions, maintained in their
own repository, standardise the span, metric and event names for model
calls, tool calls, agents and MCP. Langfuse is an open-source, self-hostable
platform that captures model and non-model calls, versions prompts, runs
evaluations (LLM judge, code, human) and tracks cost per user, built on
OpenTelemetry. Arize Phoenix is open source too, with tracing built on
OpenTelemetry and OpenInference, evaluators and datasets for comparing
versions. LangSmith and Braintrust play the same role as services. All of
them descend from Dapper (Sigelman et al., 2010), which introduced the
trace-and-span tree with propagated ids.
Go deeper
Level 2 builds the tracer in plain Python: a span, a with
block that opens and closes it, a fake clock so timings are exact, the
agent loop instrumented step by step, the export to OpenTelemetry-shaped
records, the walk that finds the first error, and the path from a
thumbs-down to a golden case. If you only needed to decide what to record
and where to send it, you are done.
Level 2
How it works, from scratch
An aeroplane's flight recorder writes down every control input, every instrument reading and every radio call, each stamped with the time. When something goes wrong, investigators don't guess; they replay the flight.
A trace is a flight recorder for one agent run. It records the whole task and every step inside it (each model call, each tool call, each search) with its inputs, outputs, timing, token counts, cost, and which model and prompt version produced it. Without traces, "why did the agent refund that order?" is guesswork. With them, it's a lookup.
Chapter 1
Spans: the steps of a run, as a tree
Everyday picture A recipe card with sub-recipes: "make the lasagne" contains "make the sauce" and "make the béchamel", each with its own start and end time.
Each step is a span: a named, timed unit of work with key-value attributes (model name, tokens, tool name…), a status (ok or error), and a parent. The top span is the whole task; everything the agent does is a child of it. All spans of one run share a trace id, so they can be gathered back together.
Worked example "Where is order A100?" produces this trace, on a clock where a model call takes 400 ms and a tool call 100 ms:
invoke_agent support-agent 900 ms [ok]
├─ chat scripted-large 400 ms [ok] in=68 out=24 tokens
├─ execute_tool lookup_order 100 ms [ok]
└─ chat scripted-large 400 ms [ok] in=127 out=14 tokens
The model asked for a tool (first chat), the tool ran, and the model
answered (second chat). Durations add up: 400 + 100 + 400 = 900 ms. The
second model call has more input tokens because the conversation, including
the tool result, is resent every step.
Figure 1 · Diagram
flowchart TD R["invoke_agent support-agent<br/>trace id tr-0001, 900 ms"] --> C1["chat: model call 1<br/>400 ms, 68 in / 24 out"] R --> T1["execute_tool lookup_order<br/>100 ms, args: A100"] R --> C2["chat: model call 2<br/>400 ms, 127 in / 14 out"]
In code: Span holds one step: its name, kind, start and end times,
attributes, status and children. walk visits a tree parents first, in time
order, and render draws it as the text tree above.
Chapter 2
Recording spans as the agent runs
Figure 2 · Diagram
sequenceDiagram participant U as User participant A as Agent loop participant T as Tracer participant M as Model participant K as Tool U->>A: Where is order A100? A->>T: open span invoke_agent A->>T: open span chat A->>M: messages + tools M-->>A: tool_use lookup_order(A100) A->>T: close span chat (tokens, model, finish reason) A->>T: open span execute_tool A->>K: lookup_order(A100) K-->>A: shipped A->>T: close span execute_tool A->>T: open span chat A->>M: messages + tool result M-->>A: "Order A100 has shipped." A->>T: close span chat A->>T: close span invoke_agent A-->>U: Order A100 has shipped.
with tracer.span(...): block, and an exception inside the
block marks the span as an error automatically.In code: Tracer.span opens a child of whatever span is open, and
closes and times it when the block ends; Span.fail marks an error that
didn't raise. FakeClock makes the timings deterministic, and
run_traced_agent is the agent loop in the diagram, recording every step.
Chapter 3
Speaking a common language: OpenTelemetry
Everyday picture Shipping containers: because every container has the same shape, any ship, crane or truck can carry any of them.
OpenTelemetry (OTel) is the open standard for traces, metrics and logs,
and its GenAI semantic conventions standardize the attribute names for
model and agent spans: gen_ai.request.model, gen_ai.usage.input_tokens,
gen_ai.tool.name and so on. Use them, and your traces flow into whatever
the organisation already runs (Datadog, Grafana, Honeycomb) or into
LLM-specific tools (Langfuse, LangSmith, Arize Phoenix, Braintrust) with no
custom glue. Attributes this module invents itself use an app. prefix,
such as app.prompt.version and app.cost_usd.
| Attribute | Meaning | Example |
|---|---|---|
gen_ai.operation.name |
what kind of step | chat, execute_tool, invoke_agent |
gen_ai.request.model |
model asked for | claude-opus-5 |
gen_ai.usage.input_tokens / output_tokens |
tokens billed | 68 / 24 |
gen_ai.response.finish_reasons |
why the model stopped | ["tool_use"] |
gen_ai.tool.name, gen_ai.tool.call.id |
which tool, which call | lookup_order, toolu_1 |
app.prompt.version |
which prompt produced this call | support-v3 |
Figure 3 · Diagram
flowchart LR A[Agent code<br/>+ tracer] -->|spans| X[Exporter] X -->|OTLP| C[Collector] C --> D[Existing monitoring<br/>Datadog, Grafana] C --> L[LLM tracing tools<br/>Langfuse, Phoenix] C --> W[Warehouse<br/>for evals and audits]
to_otel in this module produces OTLP-shaped span records.Chapter 4
Debugging a failure from its trace
Everyday picture Replaying the flight recorder to the moment the warning light came on.
Worked example A user asks "Where is order A-100?" (with a hyphen). The agent's answer is an unhelpful "I couldn't find order A-100." From the answer alone the order might simply not exist. The trace shows why in one line:
invoke_agent support-agent 900 ms [ok]
├─ chat scripted-large 400 ms [ok] in=68 out=24 tokens
├─ execute_tool lookup_order 100 ms [error] unknown order id A-100
└─ chat scripted-large 400 ms [ok] in=135 out=15 tokens
The model extracted the id exactly as typed and the tool doesn't normalize
hyphens. first_error walks the tree and returns that span. The fix
belongs in the tool (normalize ids, or return an error the model can act
on, such as "ids look like A100"), not in the prompt, and it's the trace
that tells you so.
Figure 4 · Drawn from the lesson's code
In a three-order run, four 400 ms model calls alternate back to back with three 100 ms tool calls, so model calls fill most of the 1,900 ms
Figure 5 · Drawn from the lesson's code
Input tokens climb with every model call, because each call resends the growing history
Figure 6 · Drawn from the lesson's code
In the six-order run, model calls take far more of the time than tool calls do
In code: trace_totals adds up a trace's duration, model and tool calls,
tokens, cost and errors, the numbers behind these figures.
Chapter 5
Closing the loop
Traces become most valuable when linked to feedback and evals. When a user
clicks thumbs-down, the feedback is stored against the trace id; an engineer
opens the exact trace, sees what went wrong, and turns it into a golden
test case (golden_case_from_trace, see primer.agents.evals). The fix is
made, the eval suite passes in CI (the automated checks run on every
change), and it ships.
Figure 7 · Diagram
flowchart LR P[Production traffic] --> T[Traces] T --> F[Failures flagged<br/>by users or monitors] F --> G[Add to golden set] G --> X[Fix prompt, tools<br/>or retrieval] X --> E[Evals pass in CI] E --> D[Deploy] D --> P
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1What would you record for every agent run?Think it through, then reveal
A tree of spans: the task at the root, and each model call, tool call and retrieval as children, with inputs and outputs (redacted of personal data), token counts, latency, cost, model id, prompt version, tool arguments and results, errors, and the user and tenant it ran for. That's enough to replay any decision and to aggregate cost and quality.
Question 2A user says the agent gave a wrong answer yesterday. How do you find out why?Think it through, then reveal
Look up the trace by conversation or user id and walk it: did retrieval return the right document, did the model pick the right tool with the right arguments, did a tool error, did a guardrail fire? The first error span or the first surprising output shows where the chain broke, and the trace becomes a new golden test case.
Question 3Why use OpenTelemetry instead of a vendor's SDK directly?Think it through, then reveal
Portability and fit: the organisation's existing monitoring already speaks OTel, the GenAI conventions give every tool the same attribute names, and changing vendors becomes a collector configuration change rather than a code change.
Question 4What should you be careful about when logging prompts and outputs?Think it through, then reveal
They contain user data. Redact personal data before export, restrict who can read traces, set retention limits, and keep tenants' traces separated.
Primary sources
The papers behind this lesson
Sigelman et al., Dapper, a Large-Scale Distributed Systems Tracing Infrastructure (Google technical report, 2010), Introduced the trace-and-span tree model with propagated trace ids that OpenTelemetry, and therefore agent tracing, is built on.
The paper ↗Researcher's shelf
Further reading
- OpenTelemetry semantic conventions for generative AI: https://opentelemetry.io/docs/specs/semconv/gen-ai/
- OpenTelemetry concepts: traces and spans: https://opentelemetry.io/docs/concepts/signals/traces/
- Langfuse documentation: https://langfuse.com/docs
- Arize Phoenix documentation: https://arize.com/docs/phoenix
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.