rumblr Work in progressWIP

● The AI Primer · Lesson 50 · Part 2: building systems people rely on

Cost and latency

paying less per successful task without getting worse

You'll be able to explain Routing, caching, batching, budgets, cost per successful task

Members · open during launch 24 min7 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Measure cost per successful task, not per call; include retries and human cleanup.
  2. Free wins first: prompt caching (stable prefix first), trimming tool output and chunks, concise output, batch APIs for non-interactive work.
  3. Then route easy tasks to small models, validated with evals.
  4. Run independent tool calls in parallel; stream to improve perceived speed.
  5. Budgets per task, spend caps per tenant, and alerts on sudden jumps in tokens per task.
  6. Semantic caches can serve confidently wrong answers; use them only on narrow traffic, with guards and a measured wrong-hit rate.

Level 1

The practitioner's guide

In one sentence

Cost engineering for a model-backed system means paying less per successful task, not per call, by taking the levers that can't hurt quality first and checking the ones that can against an eval.

When you need it

When the bill grows faster than the value, when a latency budget is missed, or before the first big customer arrives and the price per task stops being a rounding error. The tell: you know your spend per month but not your cost per successful task, or you know cost per call but have never counted retries and human clean-up. The lesson's worked example shows why that number, and only that number, decides things: a small model at $0.002 a call with a 60% success rate beats a large one at $0.010 and 95% if failures can simply be retried ($0.0033 against $0.0105 per success), and loses by 7x if a person has to fix each failure ($0.802 against $0.110 per task, at an illustrative $2 per fix). You don't need any of this for a prototype with ten users a day, and you shouldn't touch the levers that change behaviour (routing, semantic caching) until you have an eval to watch them with. All prices in this lesson are illustrative constants chosen for round arithmetic (a large model at $5 per million input tokens and $25 per million output, a small one at $1 and $5); check a provider's current price list for real numbers.

Your options

Eight levers, from the ones that can't hurt quality to the ones that can:

Option What it does What it guarantees What it costs Where it lives
Prompt caching Puts the stable part of the prompt first so the provider reuses its work No behaviour change; the worked 10,000-token call falls from $0.0625 to $0.0265 with 8,000 tokens cached A small write premium the first time; reordering the prompt The provider's API
Trim tokens Compress tool results, send fewer retrieved chunks, ask for concise output No change if you only cut what the next step never read; output is the expensive direction Engineering, and care not to cut what was needed Your prompt and code
Parallel tool calls Run independent tool calls at the same time Wall-clock time falls to the slowest call (300 ms to 100 ms for three lookups); tokens unchanged Concurrency in the agent loop Your agent loop
Exact response cache The same question, ignoring case and spaces, returns the stored answer Never wrong until the facts change A time-to-live to manage Your code
Batch API Non-interactive work submitted in bulk at about half price Half price, results within hours Waiting; only for work nobody is waiting on The provider's API
Budgets and alerts Caps steps and tokens per task and spend per tenant; alerts on a sudden jump A confused agent stops instead of looping all night Choosing the limits; a false alarm now and then Your code
Routing Sends easy tasks (classify, extract, format) to a small model and hard ones to a large one 53% saved on this lesson's ten-task workload, if the router is right The first lever that can lose quality; needs an eval before and after Your code, in front of every request
Semantic cache Answers a similar question from memory, using embeddings to judge similarity The cheapest possible hit Confidently wrong answers: on this lesson's pairs, 40% of different-intent questions are served from cache at a 0.95 threshold Your code, plus an embedding model

A ninth, for later: train or distil a small model on the large model's answers so it can take more of the routed traffic (Hinton et al., 2015).

How to choose

Measure first, then go down the table in order.

  • No measurement yet: instrument cost per task by step, by model, input against output and cached against uncached, and put an eval set in place to hold quality fixed.
  • A long, stable system prompt or tool list: prompt caching, today. It is the only lever that halves a bill without changing a single answer.
  • Big tool results or many retrieved chunks: trim. Compress to the fields the next step needs and send the top few reranked chunks.
  • Several independent lookups per turn: run them concurrently, and stream the answer so users see progress.
  • Anything nobody is waiting on (nightly evals, backfills, bulk classification): batch it.
  • Most traffic is simple: route, with the eval watching. A router that sends hard tasks to the small model saves money and quietly loses quality.
  • Narrow, curated, FAQ-style traffic and nothing else: a semantic cache with guards (identifiers, numbers and negation must match), a time-to-live and a measured wrong-hit rate. On everyday questions no threshold makes it safe.
  • Whatever you pick, judge it by cost per successful task, with retries and human clean-up counted in.

What it costs

The levers stack. On this lesson's ten-task workload the baseline is $0.671; caching takes it to $0.311 (2.2x), trimming tool output to $0.186 (3.6x), concise output to $0.171 (3.9x), routing to $0.086 (7.8x) and batching the non-interactive share to $0.080 (8.4x). The first three are free wins that don't touch quality; routing is the big one and the first that can. Caching is not free on the first request: Anthropic's prompt caching documentation, for one, prices a cache write at 1.25 times the base input price for a five-minute cache (2 times for an hour), reads at a tenth of it or less, and only caches prompts above a per-model minimum of a few hundred to a few thousand tokens. Batching costs time: the same provider quotes a 50% discount with most batches finishing within an hour. Parallel calls and streaming cost no tokens at all; they buy time and perceived speed. Budgets cost the occasional false alarm: this lesson's anomaly rule flags a task using more than the mean plus three standard deviations of recent usage, so a 10,000-token task against a history around 1,000 trips it at once.

What breaks

  • Cost per call replacing cost per success. The cheap model looks cheaper until failures are priced. Count retries and fixes.
  • Routing that loses quality quietly. Nothing errors; the answers just get worse. Gate the router with the eval, and re-check when the small model changes.
  • The semantic cache that answers the wrong question. "How many sick days do I get?" scores 0.97 against the vacation question and gets the vacation answer. Near-misses score higher than real paraphrases, so no threshold separates them.
  • Stale cache hits. The policy changed; the cache didn't. Give every entry a time-to-live.
  • A cache that never hits. Something volatile (a timestamp, the user's name) sits at the top of the prompt, so the prefix differs every call. Stable content first.
  • Trimming what was needed. A tool result cut to 500 tokens loses the field the next step reads. Compress by field, not by length alone.
  • The runaway loop. An agent calls the same tool with the same arguments all night. Cap steps and tokens per task; alert on tokens per task jumping.
  • One tenant's surprise invoice. A shared platform with no per-customer cap turns one customer's bug into everyone's bill.

In the wild

Prompt caching and batch processing are standard features of hosted APIs; Anthropic's documentation for both is linked in Further reading and gives the multipliers quoted above. FrugalGPT (Chen, Zaharia and Zou, 2023) named the three families, prompt adaptation, model approximation and cascades that try a cheap model first and escalate, and reported matching the best single model with up to 98% less cost on its benchmarks. Distillation (Hinton, Vinyals and Dean, 2015) is how the small model in the routing table gets good enough to take more traffic. Tooling: LiteLLM is an open-source gateway that puts one interface in front of many providers and adds routing with fallbacks, per-key and per-team budgets, spend tracking and caching; GPTCache is an open-source semantic cache built on embeddings and a vector store with a pluggable similarity evaluator, the design this lesson's SemanticCache reproduces in miniature, guards and all. Parallel tool use is part of the tool-calling protocol of the major APIs: the model asks for several tools in one turn and you return all the results in one message.

Go deeper

Level 2 prices a request symbol by symbol, builds the router, both caches and the parallel loop in plain Python, derives the anomaly alert from a mean and a standard deviation, works the unit economics with every number, and stacks the levers to draw the 8.4x figure. If you only needed to know which lever to pull first, you are done.

Level 2

How it works, from scratch

A restaurant doesn't cut costs by buying worse ingredients for every dish. It sends simple orders to the line cook and complex ones to the head chef, pre-chops what every dish shares, doesn't plate food nobody eats, cooks things in parallel, does bulk prep overnight when it's cheaper, and watches the one kitchen station whose bills suddenly spike.

Every lever in this lesson is one of those moves. The measure that matters is cost per successful task, not cost per call: a cheap attempt that fails still has to be paid for, and so does fixing it.

All prices in this module are illustrative constants chosen for round arithmetic (a "large" model at $5 per million input tokens and $25 per million output tokens; a "small" one at $1 and $5). They show the shape of the trade-offs; check your provider's current price list for real numbers.

Chapter 1

How a request is priced

Everyday picture A taxi that charges one rate for the distance to your pickup and a higher rate for the ride itself. Models charge per token (a word piece, about 4 characters of English): one rate for tokens you send (input) and a higher rate for tokens the model writes (output), because writing happens one token at a time (see primer.ml.inference).

Worked example A large-model call with 10,000 input tokens and 500 output tokens: 10,000 × $5 / 1,000,000 = $0.05 for input, plus 500 × $25 / 1,000,000 = $0.0125 for output, total $0.0625. If 8,000 of those input tokens are a stable prefix served from the prompt cache (reused work from an earlier identical beginning, billed here at 10% of the input price), input becomes 8,000 × $0.50/M + 2,000 × $5/M = $0.014, and the call costs $0.0265, 58% less.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
input tokens sent 10,000
input tokens served from the prompt cache 8,000
output tokens generated 500
price per million input / output tokens $5, $25
cached-read price as a fraction of normal input price (Greek rho) 0.1
prices are per million tokens

In words: cached input at its discounted rate, plus the rest of the input at full rate, plus output at the output rate, all per million.

On the worked example: (8,000 × 0.1 × 5 + 2,000 × 5 + 500 × 25) / 10⁶ = (4,000 + 10,000 + 12,500) / 10⁶ = $0.0265.

Level 3: in Python
n_in, c, n_out = 10_000, 8_000, 500
p_in, p_out, rho = 5, 25, 0.1
cost = (c * rho * p_in + (n_in - c) * p_in + n_out * p_out) / 10**6
round(cost, 4)  # → 0.0265

Two facts fall out: output tokens cost several times more than input (ask for concise answers), and a cached prefix is nearly free (put stable content first; see primer.agents.context). Caches usually charge a small premium the first time a prefix is written; this module ignores it for simplicity.

In code: request_cost is the formula above, reading each model's Price from PRICES.

Chapter 2

Route each task to the cheapest model that can do it

Everyday picture A hospital triage nurse: sprained ankles go to the nurse practitioner, chest pains to the cardiologist. Nobody sends every patient to the most expensive specialist.

Worked example The sample workload is 10 tasks: 7 simple (classify a ticket, extract a date) and 3 complex (plan a migration). All on the large model it costs $0.671. Routing the 7 simple ones to the small model, which is 5x cheaper per token, brings it to $0.314: 53% saved, with the hard tasks still on the strong model.

Figure 1 · Diagram

Reading it: the router is cheap (a few rules here; often a small classifier) and sits in front of every request. The whole saving comes from the left branch: most traffic in real systems is simple, and simple work runs well on small models. Validate the router with evals (primer.agents.evals): a router that sends hard tasks to the small model saves money and quietly loses quality.

In code: route is the rule-based router; workload_cost prices a list of WorkItem tasks with any levers switched on; routing_savings compares all-large against routed on SAMPLE_WORKLOAD.

Chapter 3

Response caching and semantic caching

Everyday picture A receptionist who's been asked "what's the wifi password?" a hundred times just answers from memory. That's an exact response cache: the same question (ignoring case and spaces) returns the stored answer with no model call. A semantic cache goes further and answers similar questions from memory, using embeddings (vectors where closeness means similar meaning, see primer.ml.embeddings) to decide "similar". That's where the receptionist starts giving the vacation-policy answer to someone asking about sick leave.

Worked example With the toy embedder in primer.common.embedder:

Stored question New question Similarity Same intent?
How do I reset my password? I forgot my password, how do I recover it? 0.88 yes
How many vacation days do I get? How many sick days do I get? 0.97 no
Request a new laptop Request a new monitor 0.96 no
What does ERR-4012 mean? What does ERR-4013 mean? 0.56 no

The two near-misses score higher than a genuine paraphrase. No single threshold separates them. Two cheap guards help: identifiers and numbers must match exactly (ERR-4012 ≠ ERR-4013), and negation must match ("cancel" ≠ "don't cancel"). They can't catch "sick" vs "vacation", which share a topic and differ in a single word.

Figure 3 · Diagram

Reading it: the exact cache is checked first because it's free and never wrong. The semantic path adds two gates: the similarity threshold, and the guards. An expired entry (older than its TTL, time to live) is treated as a miss, so answers about changing facts don't go stale forever.

Figure 2 · Drawn from the lesson's code

0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 similarity threshold 0.0 0.2 0.4 0.6 0.8 share of pairs Semantic cache: no threshold separates them cleanly paraphrases served from cache (good) different questions served from cache (wrong)

No threshold separates them: wrong hits stay at 40 to 60% through 0.95, often above the paraphrase hit rate, and reach 0 only where paraphrase hits do too

Reading it: the x-axis is the similarity threshold; the blue line is the share of genuine paraphrases answered from cache (good), and the red line is the share of different-intent questions answered from cache (a wrong answer served confidently). Lowering the threshold raises both, but look where the lines sit: from 0.4 upward the red line is above the blue one, so the cache serves more wrong answers than right ones (60% of different-intent questions against at most 57% of paraphrases). Even at 0.95 it still answers 40% of them wrongly and only 14% of paraphrases, because "sick" vs "vacation" scores 0.97. On everyday questions no threshold makes semantic caching safe; it's safe only for narrow, curated FAQ-style traffic, with guards, a TTL, and a measured wrong-hit rate.

In code: ResponseCache is the exact cache. SemanticCache is the semantic path: SemanticCache.lookup skips expired entries and entries whose identifiers or negation differ, then returns the nearest survivor, and SemanticCache.get applies the threshold. semantic_cache_sweep draws the figure from the labelled pairs in CACHE_PAIRS.

Chapter 4

Trim tokens

Everyday picture Don't photocopy the whole binder when the colleague needs one page. Compress tool results to the fields the next step needs (primer.agents.context.compress_tool_output), send the top few reranked chunks instead of dozens, remove repeated boilerplate from system prompts, and ask for concise output, because output is the expensive direction.

In code: workload_cost models trimming as cutting each task's tool output to 500 tokens, and concise output as cutting long answers by a third.

Chapter 5

Run independent tool calls at the same time

Everyday picture Boil the pasta while the sauce simmers. Cooking them one after the other takes the sum of the times; together, the time of the slowest.

Worked example Three independent lookups of 100 ms each: sequentially about 300 ms; concurrently about 100 ms.

Figure 5 · Diagram

Reading it: in the top half each call waits for the previous result; in the bottom half the three requests leave together and the agent waits only for the slowest. Models can ask for several tools in one turn (parallel tool use); run them concurrently and return all results in one message. Parallel calls cut wall-clock time, not tokens.

Figure 4 · Drawn from the lesson's code

0 50 100 150 200 250 300 milliseconds since the agent started the calls sequential parallel Three independent 100 ms tool calls tool 0 tool 1 tool 2 tool 0 tool 1 tool 2

Three 100 ms tool calls take 300 ms end to end when run one after another, but only 100 ms when started together

Reading it: each bar is one 100 ms tool call on a shared time axis. The sequential calls stack end to end, 300 ms in all; the parallel ones start together, so the whole batch finishes in the time of one. Run the lesson to measure it with asyncio: the real timings land within a few milliseconds of these bars.

In code: run_sequential awaits each call before starting the next; run_parallel starts them all with Python's asyncio gather and waits once.

Streaming (showing tokens as they're generated) doesn't reduce total time either, but users see progress immediately, which changes how fast the system feels.

Chapter 6

Batch what isn't interactive

Everyday picture Sending the laundry out to be done overnight at half price instead of waiting at the express counter. Providers offer batch APIs: submit many requests, get results within hours (often within 24), typically at about half the price. Use them for anything no one is waiting on: nightly evals, backfilling document processing, bulk classification.

In code: batch_cost prices a list of requests at the batch discount.

Chapter 7

Budgets and alerts

Everyday picture A prepaid card with a hard limit, and a bank that texts you when a purchase looks nothing like your usual spending.

A task budget caps steps and tokens per task, so a confused agent stops instead of looping all night. A tenant spend cap (a tenant is one customer organisation on a shared platform) stops one customer's runaway usage from becoming a surprise invoice. An anomaly alert fires when a task uses far more tokens than usual:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
tokens used by the task just finished 10,000
mean tokens per task over this tenant's recent history (Greek mu) 1,000
standard deviation of that history: the typical distance from the mean (Greek sigma) ≈ 71
how many standard deviations count as unusual 3

In words: alert when a task uses more tokens than the usual amount plus three times the usual spread.

On the worked example: history 1000, 1100, 900, 1050, 950, 1000 has mean 1,000 and standard deviation ≈ 71 (the squared distances from the mean add up to 25,000; divided by 6 − 1 = 5 that's 5,000, whose square root is 70.7), so the line is 1,000 + 3 × 70.7 ≈ 1,212. A 10,000-token task is far above it: alert. Such jumps usually mean a loop or a bad deploy.

Level 3: in Python
import statistics
history = [1000, 1100, 900, 1050, 950, 1000]
mu = statistics.mean(history)
# divides by 6 - 1, as above
sigma = statistics.stdev(history)
mu, round(sigma, 1)  # → (1000, 70.7)
z = 3
# the alert line
round(mu + z * sigma)  # → 1212
# alert?
10_000 > mu + z * sigma  # → True

In code: TaskBudget.charge counts each step's tokens and raises BudgetExceeded at either limit; TenantSpend.record adds a finished task's cost to its tenant's total and returns a cap alert or an anomaly alert (the formula above).

Chapter 8

Unit economics: cost per successful task

Everyday picture A cheap printer that jams on 40% of pages isn't cheap once you count the wasted paper and the time spent clearing jams.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Small model Large model
model cost of one attempt $0.002 $0.010
chance an attempt succeeds 0.60 0.95
expected number of attempts until one succeeds 1.67 1.05
cost of a person fixing one failure (illustrative: a few minutes of staff time) $2.00 $2.00

In words: if you can simply retry, you pay one attempt's cost for every expected attempt; if failures need a person, every task pays its attempt plus the chance of failure times the cost of the fix.

On the worked example: retrying, the small model costs 0.002 / 0.60 = $0.0033 per success and the large one 0.010 / 0.95 = $0.0105, so the small model wins. With human cleanup, the small model costs 0.002 + 0.40 × 2.00 = $0.802 per task and the large one 0.010 + 0.05 × 2.00 = $0.110: the "cheap" model is 7x more expensive.

Level 3: in Python
def per_success(c, p):
    # retry until it works
    return c / p
def per_task(c, p, h=2.00):
    # a person fixes each failure
    return c + (1 - p) * h
round(per_success(0.002, 0.60), 4), round(per_success(0.010, 0.95), 4)  # → (0.0033, 0.0105)
round(per_task(0.002, 0.60), 3), round(per_task(0.010, 0.95), 3)  # → (0.802, 0.11)
# small vs. large, with cleanup
round(per_task(0.002, 0.60) / per_task(0.010, 0.95))  # → 7

Figure 6 · Drawn from the lesson's code

retry until success person fixes failures 1 0 − 2 1 0 − 1 1 0 0 dollars per successful task (log scale) The cheap model is only cheap if failures are free $0.003 $0.011 $0.802 $0.110 small: $0.002/attempt, 60% success large: $0.010/attempt, 95% success

With automatic retries the small model is cheaper per success (0.3 vs 1.1 cents); if a person fixes each failure the large model wins, 11 cents vs 80

Reading it: the left pair of bars assumes failures can be retried automatically: the small model wins. The right pair assumes a person has to fix each failure: the large model wins by a mile. Which world you're in decides which model is cheaper, and the model's price per token barely matters in the second one.

In code: cost_per_success_with_retries is and cost_per_success_with_cleanup is .

Chapter 9

Putting it together: a 5x plan, in order

Order the levers so the ones that can't hurt quality come first:

  1. Prompt caching: reorder the prompt so the stable prefix is reused. No behaviour change.
  2. Trim tokens: compress tool results and send fewer chunks.
  3. Concise output: ask for shorter answers where length adds nothing.
  4. Route easy tasks to the small model: the first lever that can change quality, so it's validated with evals before and after.
  5. Batch the non-interactive share.

Figure 7 · Drawn from the lesson's code

baseline + caching + trim + concise + routing + batch 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 workload cost ($) Levers applied in order (blue: can't change quality; orange: validate with evals) 1.0x 2.2x 3.6x 3.9x 7.8x 8.4x

Workload cost falls from 67 to 8 cents as levers stack: caching 2.2x, trimming 3.6x, concise output 3.9x, routing 7.8x, batching 8.4x

Reading it: each bar is the cost of the same 10-task workload after applying every lever up to and including that one; the label is the cumulative reduction. Caching alone roughly halves it, trimming and concise output take it past 3x, routing takes it past 7x, and batching adds the last few percent. The three levers after the baseline (caching, trimming, concise output) are the free wins that don't touch quality.

In code: five_x_plan switches the levers on one at a time, in this order, and reports the cost and cumulative reduction after each.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1How would you cut the cost per task by 5x without hurting quality?Think it through, then reveal

First measure: cost per successful task, broken down by step, model, input vs. output, and cached vs. uncached tokens, with an eval set to hold quality fixed. Then, in order: restructure prompts for prompt caching; compress tool results and send only the top reranked chunks; ask for concise output; route simple steps (classification, extraction, formatting) to a small model, checking the eval before and after; move anything non-interactive to a batch API; cap steps and tokens per task. On the sample workload here that's about 8x, and every step after the first three is checked against the eval set.

Question 2Why can a cheaper model cost more?Think it through, then reveal

Because failures cost money too: retries, human cleanup, lost customers. At 60% success with a $2 human fix, a $0.002 call costs $0.80 per task; a $0.010 call at 95% costs $0.11.

Question 3What are the risks of a semantic cache?Think it through, then reveal

Returning a confident, wrong answer to a question that's similar but not the same ("sick days" vs "vacation days"), and serving stale answers after facts change. Mitigate with high thresholds, exact-match guards on IDs, numbers and negation, per-intent scoping, a TTL, and a measured wrong-hit rate on labelled pairs before enabling it.

Question 4Your average tokens per task doubled overnight. What do you check?Think it through, then reveal

Whether a deploy changed a prompt or a tool (a bigger tool output, a new retrieval setting), whether an agent is looping (the same tool called with the same arguments), and whether the prompt cache hit rate dropped (something volatile moved to the top). Traces make this a lookup rather than a guess (primer.agents.observability).

Primary sources

The papers behind this lesson

Hinton, Vinyals & Dean, Distilling the Knowledge in a Neural Network (2015), Showed how to train a small model to imitate a large one, the technique behind making the cheap model good enough to take more of the routed traffic.

Read the annotated companion →The paper ↗

Chen, Zaharia & Zou, FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance (2023), Studied prompt adaptation, caching and cascades that try cheap models first and escalate to expensive ones only when needed.

The paper ↗

Researcher's shelf

Further reading

  • Anthropic prompt caching docs: https://docs.claude.com/en/docs/build-with-claude/prompt-caching
  • Anthropic Message Batches docs: https://docs.claude.com/en/docs/build-with-claude/batch-processing
  • Anthropic, parallel tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
  • Python asyncio.gather: https://docs.python.org/3/library/asyncio-task.html#asyncio.gather

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.