The lesson in one minute
What you'll be able to explain
- Measure cost per successful task, not per call; include retries and human cleanup.
- Free wins first: prompt caching (stable prefix first), trimming tool output and chunks, concise output, batch APIs for non-interactive work.
- Then route easy tasks to small models, validated with evals.
- Run independent tool calls in parallel; stream to improve perceived speed.
- Budgets per task, spend caps per tenant, and alerts on sudden jumps in tokens per task.
- Semantic caches can serve confidently wrong answers; use them only on narrow traffic, with guards and a measured wrong-hit rate.
Level 1
The practitioner's guide
In one sentence
Cost engineering for a model-backed system means paying less per successful task, not per call, by taking the levers that can't hurt quality first and checking the ones that can against an eval.
When you need it
When the bill grows faster than the value, when a latency budget is missed, or before the first big customer arrives and the price per task stops being a rounding error. The tell: you know your spend per month but not your cost per successful task, or you know cost per call but have never counted retries and human clean-up. The lesson's worked example shows why that number, and only that number, decides things: a small model at $0.002 a call with a 60% success rate beats a large one at $0.010 and 95% if failures can simply be retried ($0.0033 against $0.0105 per success), and loses by 7x if a person has to fix each failure ($0.802 against $0.110 per task, at an illustrative $2 per fix). You don't need any of this for a prototype with ten users a day, and you shouldn't touch the levers that change behaviour (routing, semantic caching) until you have an eval to watch them with. All prices in this lesson are illustrative constants chosen for round arithmetic (a large model at $5 per million input tokens and $25 per million output, a small one at $1 and $5); check a provider's current price list for real numbers.
Your options
Eight levers, from the ones that can't hurt quality to the ones that can:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Prompt caching | Puts the stable part of the prompt first so the provider reuses its work | No behaviour change; the worked 10,000-token call falls from $0.0625 to $0.0265 with 8,000 tokens cached | A small write premium the first time; reordering the prompt | The provider's API |
| Trim tokens | Compress tool results, send fewer retrieved chunks, ask for concise output | No change if you only cut what the next step never read; output is the expensive direction | Engineering, and care not to cut what was needed | Your prompt and code |
| Parallel tool calls | Run independent tool calls at the same time | Wall-clock time falls to the slowest call (300 ms to 100 ms for three lookups); tokens unchanged | Concurrency in the agent loop | Your agent loop |
| Exact response cache | The same question, ignoring case and spaces, returns the stored answer | Never wrong until the facts change | A time-to-live to manage | Your code |
| Batch API | Non-interactive work submitted in bulk at about half price | Half price, results within hours | Waiting; only for work nobody is waiting on | The provider's API |
| Budgets and alerts | Caps steps and tokens per task and spend per tenant; alerts on a sudden jump | A confused agent stops instead of looping all night | Choosing the limits; a false alarm now and then | Your code |
| Routing | Sends easy tasks (classify, extract, format) to a small model and hard ones to a large one | 53% saved on this lesson's ten-task workload, if the router is right | The first lever that can lose quality; needs an eval before and after | Your code, in front of every request |
| Semantic cache | Answers a similar question from memory, using embeddings to judge similarity | The cheapest possible hit | Confidently wrong answers: on this lesson's pairs, 40% of different-intent questions are served from cache at a 0.95 threshold | Your code, plus an embedding model |
A ninth, for later: train or distil a small model on the large model's answers so it can take more of the routed traffic (Hinton et al., 2015).
How to choose
Measure first, then go down the table in order.
- No measurement yet: instrument cost per task by step, by model, input against output and cached against uncached, and put an eval set in place to hold quality fixed.
- A long, stable system prompt or tool list: prompt caching, today. It is the only lever that halves a bill without changing a single answer.
- Big tool results or many retrieved chunks: trim. Compress to the fields the next step needs and send the top few reranked chunks.
- Several independent lookups per turn: run them concurrently, and stream the answer so users see progress.
- Anything nobody is waiting on (nightly evals, backfills, bulk classification): batch it.
- Most traffic is simple: route, with the eval watching. A router that sends hard tasks to the small model saves money and quietly loses quality.
- Narrow, curated, FAQ-style traffic and nothing else: a semantic cache with guards (identifiers, numbers and negation must match), a time-to-live and a measured wrong-hit rate. On everyday questions no threshold makes it safe.
- Whatever you pick, judge it by cost per successful task, with retries and human clean-up counted in.
What it costs
The levers stack. On this lesson's ten-task workload the baseline is $0.671; caching takes it to $0.311 (2.2x), trimming tool output to $0.186 (3.6x), concise output to $0.171 (3.9x), routing to $0.086 (7.8x) and batching the non-interactive share to $0.080 (8.4x). The first three are free wins that don't touch quality; routing is the big one and the first that can. Caching is not free on the first request: Anthropic's prompt caching documentation, for one, prices a cache write at 1.25 times the base input price for a five-minute cache (2 times for an hour), reads at a tenth of it or less, and only caches prompts above a per-model minimum of a few hundred to a few thousand tokens. Batching costs time: the same provider quotes a 50% discount with most batches finishing within an hour. Parallel calls and streaming cost no tokens at all; they buy time and perceived speed. Budgets cost the occasional false alarm: this lesson's anomaly rule flags a task using more than the mean plus three standard deviations of recent usage, so a 10,000-token task against a history around 1,000 trips it at once.
What breaks
- Cost per call replacing cost per success. The cheap model looks cheaper until failures are priced. Count retries and fixes.
- Routing that loses quality quietly. Nothing errors; the answers just get worse. Gate the router with the eval, and re-check when the small model changes.
- The semantic cache that answers the wrong question. "How many sick days do I get?" scores 0.97 against the vacation question and gets the vacation answer. Near-misses score higher than real paraphrases, so no threshold separates them.
- Stale cache hits. The policy changed; the cache didn't. Give every entry a time-to-live.
- A cache that never hits. Something volatile (a timestamp, the user's name) sits at the top of the prompt, so the prefix differs every call. Stable content first.
- Trimming what was needed. A tool result cut to 500 tokens loses the field the next step reads. Compress by field, not by length alone.
- The runaway loop. An agent calls the same tool with the same arguments all night. Cap steps and tokens per task; alert on tokens per task jumping.
- One tenant's surprise invoice. A shared platform with no per-customer cap turns one customer's bug into everyone's bill.
In the wild
Prompt caching and batch processing are standard features
of hosted APIs; Anthropic's documentation for both is linked in Further
reading and gives the multipliers quoted above. FrugalGPT (Chen, Zaharia and
Zou, 2023) named the three families, prompt adaptation, model approximation
and cascades that try a cheap model first and escalate, and reported
matching the best single model with up to 98% less cost on its benchmarks.
Distillation (Hinton, Vinyals and Dean, 2015) is how the small model in the
routing table gets good enough to take more traffic. Tooling: LiteLLM is an
open-source gateway that puts one interface in front of many providers and
adds routing with fallbacks, per-key and per-team budgets, spend tracking
and caching; GPTCache is an open-source semantic cache built on embeddings
and a vector store with a pluggable similarity evaluator, the design this
lesson's SemanticCache reproduces in miniature, guards and all. Parallel
tool use is part of the tool-calling protocol of the major APIs: the model
asks for several tools in one turn and you return all the results in one
message.
Go deeper
Level 2 prices a request symbol by symbol, builds the router, both caches and the parallel loop in plain Python, derives the anomaly alert from a mean and a standard deviation, works the unit economics with every number, and stacks the levers to draw the 8.4x figure. If you only needed to know which lever to pull first, you are done.
Level 2
How it works, from scratch
A restaurant doesn't cut costs by buying worse ingredients for every dish. It sends simple orders to the line cook and complex ones to the head chef, pre-chops what every dish shares, doesn't plate food nobody eats, cooks things in parallel, does bulk prep overnight when it's cheaper, and watches the one kitchen station whose bills suddenly spike.
Every lever in this lesson is one of those moves. The measure that matters is cost per successful task, not cost per call: a cheap attempt that fails still has to be paid for, and so does fixing it.
All prices in this module are illustrative constants chosen for round arithmetic (a "large" model at $5 per million input tokens and $25 per million output tokens; a "small" one at $1 and $5). They show the shape of the trade-offs; check your provider's current price list for real numbers.
Chapter 1
How a request is priced
Everyday picture A taxi that charges one rate for the distance to your
pickup and a higher rate for the ride itself. Models charge per token
(a word piece, about 4 characters of English): one rate for tokens you send
(input) and a higher rate for tokens the model writes (output),
because writing happens one token at a time (see primer.ml.inference).
Worked example A large-model call with 10,000 input tokens and 500 output tokens: 10,000 × $5 / 1,000,000 = $0.05 for input, plus 500 × $25 / 1,000,000 = $0.0125 for output, total $0.0625. If 8,000 of those input tokens are a stable prefix served from the prompt cache (reused work from an earlier identical beginning, billed here at 10% of the input price), input becomes 8,000 × $0.50/M + 2,000 × $5/M = $0.014, and the call costs $0.0265, 58% less.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| input tokens sent | 10,000 | |
| input tokens served from the prompt cache | 8,000 | |
| output tokens generated | 500 | |
| price per million input / output tokens | $5, $25 | |
| cached-read price as a fraction of normal input price (Greek rho) | 0.1 | |
| prices are per million tokens |
In words: cached input at its discounted rate, plus the rest of the input at full rate, plus output at the output rate, all per million.
On the worked example: (8,000 × 0.1 × 5 + 2,000 × 5 + 500 × 25) / 10⁶ = (4,000 + 10,000 + 12,500) / 10⁶ = $0.0265.
Level 3: in Python
n_in, c, n_out = 10_000, 8_000, 500
p_in, p_out, rho = 5, 25, 0.1
cost = (c * rho * p_in + (n_in - c) * p_in + n_out * p_out) / 10**6
round(cost, 4) # → 0.0265
Two facts fall out: output tokens cost several times more than input (ask
for concise answers), and a cached prefix is nearly free (put stable content
first; see primer.agents.context). Caches usually charge a small premium
the first time a prefix is written; this module ignores it for simplicity.
In code: request_cost is the formula above, reading each model's
Price from PRICES.
Chapter 2
Route each task to the cheapest model that can do it
Everyday picture A hospital triage nurse: sprained ankles go to the nurse practitioner, chest pains to the cardiologist. Nobody sends every patient to the most expensive specialist.
Worked example The sample workload is 10 tasks: 7 simple (classify a ticket, extract a date) and 3 complex (plan a migration). All on the large model it costs $0.671. Routing the 7 simple ones to the small model, which is 5x cheaper per token, brings it to $0.314: 53% saved, with the hard tasks still on the strong model.
Figure 1 · Diagram
flowchart LR
T[Incoming task] --> R{Router<br/>rules or a small classifier}
R -->|classify, extract,<br/>format, short| S[Small, fast model]
R -->|plan, analyze,<br/>multi-step reasoning| L[Large model]
S --> O[Result]
L --> O
primer.agents.evals): a router that sends hard tasks to the small model
saves money and quietly loses quality.In code: route is the rule-based router; workload_cost prices a list
of WorkItem tasks with any levers switched on; routing_savings compares
all-large against routed on SAMPLE_WORKLOAD.
Chapter 3
Response caching and semantic caching
Everyday picture A receptionist who's been asked "what's the wifi
password?" a hundred times just answers from memory. That's an exact
response cache: the same question (ignoring case and spaces) returns the
stored answer with no model call. A semantic cache goes further and
answers similar questions from memory, using embeddings (vectors where
closeness means similar meaning, see primer.ml.embeddings) to decide
"similar". That's where the receptionist starts giving the vacation-policy
answer to someone asking about sick leave.
Worked example With the toy embedder in primer.common.embedder:
| Stored question | New question | Similarity | Same intent? |
|---|---|---|---|
| How do I reset my password? | I forgot my password, how do I recover it? | 0.88 | yes |
| How many vacation days do I get? | How many sick days do I get? | 0.97 | no |
| Request a new laptop | Request a new monitor | 0.96 | no |
| What does ERR-4012 mean? | What does ERR-4013 mean? | 0.56 | no |
The two near-misses score higher than a genuine paraphrase. No single threshold separates them. Two cheap guards help: identifiers and numbers must match exactly (ERR-4012 ≠ ERR-4013), and negation must match ("cancel" ≠ "don't cancel"). They can't catch "sick" vs "vacation", which share a topic and differ in a single word.
Figure 3 · Diagram
flowchart TD
Q[New question] --> E{Exact match<br/>in response cache?}
E -->|yes| A1[Return stored answer]
E -->|no| V[Embed and find the<br/>nearest stored question]
V --> T{Similarity above<br/>threshold?}
T -->|no| M[Call the model]
T -->|yes| G{IDs, numbers and<br/>negation identical?<br/>entry not expired?}
G -->|no| M
G -->|yes| A2[Return stored answer]
M --> S[Store the new answer]
Figure 2 · Drawn from the lesson's code
No threshold separates them: wrong hits stay at 40 to 60% through 0.95, often above the paraphrase hit rate, and reach 0 only where paraphrase hits do too
In code: ResponseCache is the exact cache. SemanticCache is the
semantic path: SemanticCache.lookup skips expired entries and entries whose
identifiers or negation differ, then returns the nearest survivor, and
SemanticCache.get applies the threshold. semantic_cache_sweep draws the
figure from the labelled pairs in CACHE_PAIRS.
Chapter 4
Trim tokens
Everyday picture Don't photocopy the whole binder when the colleague
needs one page. Compress tool results to the fields the next step needs
(primer.agents.context.compress_tool_output), send the top few reranked
chunks instead of dozens, remove repeated boilerplate from system prompts,
and ask for concise output, because output is the expensive direction.
In code: workload_cost models trimming as cutting each task's tool
output to 500 tokens, and concise output as cutting long answers by a third.
Chapter 5
Run independent tool calls at the same time
Everyday picture Boil the pasta while the sauce simmers. Cooking them one after the other takes the sum of the times; together, the time of the slowest.
Worked example Three independent lookups of 100 ms each: sequentially about 300 ms; concurrently about 100 ms.
Figure 5 · Diagram
sequenceDiagram
participant A as Agent
participant T1 as Weather API
participant T2 as Calendar API
participant T3 as CRM API
Note over A,T3: Sequential: about 300 ms
A->>T1: call
T1-->>A: result
A->>T2: call
T2-->>A: result
A->>T3: call
T3-->>A: result
Note over A,T3: Parallel: about 100 ms
par
A->>T1: call
and
A->>T2: call
and
A->>T3: call
end
T1-->>A: result
T2-->>A: result
T3-->>A: result
Figure 4 · Drawn from the lesson's code
Three 100 ms tool calls take 300 ms end to end when run one after another, but only 100 ms when started together
asyncio: the real timings land within a few milliseconds
of these bars.In code: run_sequential awaits each call before starting the next;
run_parallel starts them all with Python's asyncio gather and waits once.
Streaming (showing tokens as they're generated) doesn't reduce total time either, but users see progress immediately, which changes how fast the system feels.
Chapter 6
Batch what isn't interactive
Everyday picture Sending the laundry out to be done overnight at half price instead of waiting at the express counter. Providers offer batch APIs: submit many requests, get results within hours (often within 24), typically at about half the price. Use them for anything no one is waiting on: nightly evals, backfilling document processing, bulk classification.
In code: batch_cost prices a list of requests at the batch discount.
Chapter 7
Budgets and alerts
Everyday picture A prepaid card with a hard limit, and a bank that texts you when a purchase looks nothing like your usual spending.
A task budget caps steps and tokens per task, so a confused agent stops instead of looping all night. A tenant spend cap (a tenant is one customer organisation on a shared platform) stops one customer's runaway usage from becoming a surprise invoice. An anomaly alert fires when a task uses far more tokens than usual:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Worked example |
|---|---|---|
| tokens used by the task just finished | 10,000 | |
| mean tokens per task over this tenant's recent history (Greek mu) | 1,000 | |
| standard deviation of that history: the typical distance from the mean (Greek sigma) | ≈ 71 | |
| how many standard deviations count as unusual | 3 |
In words: alert when a task uses more tokens than the usual amount plus three times the usual spread.
On the worked example: history 1000, 1100, 900, 1050, 950, 1000 has mean 1,000 and standard deviation ≈ 71 (the squared distances from the mean add up to 25,000; divided by 6 − 1 = 5 that's 5,000, whose square root is 70.7), so the line is 1,000 + 3 × 70.7 ≈ 1,212. A 10,000-token task is far above it: alert. Such jumps usually mean a loop or a bad deploy.
Level 3: in Python
import statistics
history = [1000, 1100, 900, 1050, 950, 1000]
mu = statistics.mean(history)
# divides by 6 - 1, as above
sigma = statistics.stdev(history)
mu, round(sigma, 1) # → (1000, 70.7)
z = 3
# the alert line
round(mu + z * sigma) # → 1212
# alert?
10_000 > mu + z * sigma # → True
In code: TaskBudget.charge counts each step's tokens and raises
BudgetExceeded at either limit; TenantSpend.record adds a finished task's
cost to its tenant's total and returns a cap alert or an anomaly alert (the
formula above).
Chapter 8
Unit economics: cost per successful task
Everyday picture A cheap printer that jams on 40% of pages isn't cheap once you count the wasted paper and the time spent clearing jams.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Small model | Large model |
|---|---|---|---|
| model cost of one attempt | $0.002 | $0.010 | |
| chance an attempt succeeds | 0.60 | 0.95 | |
| expected number of attempts until one succeeds | 1.67 | 1.05 | |
| cost of a person fixing one failure (illustrative: a few minutes of staff time) | $2.00 | $2.00 |
In words: if you can simply retry, you pay one attempt's cost for every expected attempt; if failures need a person, every task pays its attempt plus the chance of failure times the cost of the fix.
On the worked example: retrying, the small model costs 0.002 / 0.60 = $0.0033 per success and the large one 0.010 / 0.95 = $0.0105, so the small model wins. With human cleanup, the small model costs 0.002 + 0.40 × 2.00 = $0.802 per task and the large one 0.010 + 0.05 × 2.00 = $0.110: the "cheap" model is 7x more expensive.
Level 3: in Python
def per_success(c, p):
# retry until it works
return c / p
def per_task(c, p, h=2.00):
# a person fixes each failure
return c + (1 - p) * h
round(per_success(0.002, 0.60), 4), round(per_success(0.010, 0.95), 4) # → (0.0033, 0.0105)
round(per_task(0.002, 0.60), 3), round(per_task(0.010, 0.95), 3) # → (0.802, 0.11)
# small vs. large, with cleanup
round(per_task(0.002, 0.60) / per_task(0.010, 0.95)) # → 7
Figure 6 · Drawn from the lesson's code
With automatic retries the small model is cheaper per success (0.3 vs 1.1 cents); if a person fixes each failure the large model wins, 11 cents vs 80
In code: cost_per_success_with_retries is and
cost_per_success_with_cleanup is .
Chapter 9
Putting it together: a 5x plan, in order
Order the levers so the ones that can't hurt quality come first:
- Prompt caching: reorder the prompt so the stable prefix is reused. No behaviour change.
- Trim tokens: compress tool results and send fewer chunks.
- Concise output: ask for shorter answers where length adds nothing.
- Route easy tasks to the small model: the first lever that can change quality, so it's validated with evals before and after.
- Batch the non-interactive share.
Figure 7 · Drawn from the lesson's code
Workload cost falls from 67 to 8 cents as levers stack: caching 2.2x, trimming 3.6x, concise output 3.9x, routing 7.8x, batching 8.4x
In code: five_x_plan switches the levers on one at a time, in this
order, and reports the cost and cumulative reduction after each.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1How would you cut the cost per task by 5x without hurting quality?Think it through, then reveal
First measure: cost per successful task, broken down by step, model, input vs. output, and cached vs. uncached tokens, with an eval set to hold quality fixed. Then, in order: restructure prompts for prompt caching; compress tool results and send only the top reranked chunks; ask for concise output; route simple steps (classification, extraction, formatting) to a small model, checking the eval before and after; move anything non-interactive to a batch API; cap steps and tokens per task. On the sample workload here that's about 8x, and every step after the first three is checked against the eval set.
Question 2Why can a cheaper model cost more?Think it through, then reveal
Because failures cost money too: retries, human cleanup, lost customers. At 60% success with a $2 human fix, a $0.002 call costs $0.80 per task; a $0.010 call at 95% costs $0.11.
Question 3What are the risks of a semantic cache?Think it through, then reveal
Returning a confident, wrong answer to a question that's similar but not the same ("sick days" vs "vacation days"), and serving stale answers after facts change. Mitigate with high thresholds, exact-match guards on IDs, numbers and negation, per-intent scoping, a TTL, and a measured wrong-hit rate on labelled pairs before enabling it.
Question 4Your average tokens per task doubled overnight. What do you check?Think it through, then reveal
Whether a deploy changed a prompt or a tool (a bigger tool output, a new
retrieval setting), whether an agent is looping (the same tool called with
the same arguments), and whether the prompt cache hit rate dropped
(something volatile moved to the top). Traces make this a lookup rather
than a guess (primer.agents.observability).
Primary sources
The papers behind this lesson
Hinton, Vinyals & Dean, Distilling the Knowledge in a Neural Network (2015), Showed how to train a small model to imitate a large one, the technique behind making the cheap model good enough to take more of the routed traffic.
Read the annotated companion →The paper ↗Chen, Zaharia & Zou, FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance (2023), Studied prompt adaptation, caching and cascades that try cheap models first and escalate to expensive ones only when needed.
The paper ↗Researcher's shelf
Further reading
- Anthropic prompt caching docs: https://docs.claude.com/en/docs/build-with-claude/prompt-caching
- Anthropic Message Batches docs: https://docs.claude.com/en/docs/build-with-claude/batch-processing
- Anthropic, parallel tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
- Python
asyncio.gather: https://docs.python.org/3/library/asyncio-task.html#asyncio.gather
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.