rumblr Work in progressWIP

● The AI Primer · Lesson 50 · Part 2: building systems people rely on

Cost and latency

paying less per successful task without getting worse

This lesson covers Routing, caching, batching, budgets, cost per successful task

Members · open during launch 24 min7 figures and diagrams
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Measure cost per successful task, not per call; include retries and human cleanup.
  2. Free wins first: prompt caching (stable prefix first), trimming tool output and chunks, concise output, batch APIs for non-interactive work.
  3. Then route easy tasks to small models, validated with evals.
  4. Run independent tool calls in parallel; stream to improve perceived speed.
  5. Budgets per task, spend caps per tenant, and alerts on sudden jumps in tokens per task.
  6. Semantic caches can serve confidently wrong answers; use them only on narrow traffic, with guards and a measured wrong-hit rate.

Level 2

How it works, from scratch

A restaurant doesn't cut costs by buying worse ingredients for every dish. It sends simple orders to the line cook and complex ones to the head chef, pre-chops what every dish shares, doesn't plate food nobody eats, cooks things in parallel, does bulk prep overnight when it's cheaper, and watches the one kitchen station whose bills suddenly spike.

Every lever in this lesson is one of those moves. The measure that matters is cost per successful task, not cost per call: a cheap attempt that fails still has to be paid for, and so does fixing it.

All prices in this module are illustrative constants chosen for round arithmetic (a "large" model at $5 per million input tokens and $25 per million output tokens; a "small" one at $1 and $5). They show the shape of the trade-offs; check your provider's current price list for real numbers.

Chapter 1

How a request is priced

Everyday picture A taxi that charges one rate for the distance to your pickup and a higher rate for the ride itself. Models charge per token (a word piece, about 4 characters of English): one rate for tokens you send (input) and a higher rate for tokens the model writes (output), because writing happens one token at a time (see primer.ml.inference).

Worked example A large-model call with 10,000 input tokens and 500 output tokens: 10,000 × $5 / 1,000,000 = $0.05 for input, plus 500 × $25 / 1,000,000 = $0.0125 for output, total $0.0625. If 8,000 of those input tokens are a stable prefix served from the prompt cache (reused work from an earlier identical beginning, billed here at 10% of the input price), input becomes 8,000 × $0.50/M + 2,000 × $5/M = $0.014, and the call costs $0.0265, 58% less.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
input tokens sent 10,000
input tokens served from the prompt cache 8,000
output tokens generated 500
price per million input / output tokens $5, $25
cached-read price as a fraction of normal input price (Greek rho) 0.1
prices are per million tokens

In words: cached input at its discounted rate, plus the rest of the input at full rate, plus output at the output rate, all per million.

On the worked example: (8,000 × 0.1 × 5 + 2,000 × 5 + 500 × 25) / 10⁶ = (4,000 + 10,000 + 12,500) / 10⁶ = $0.0265.

Level 3: in Python
n_in, c, n_out = 10_000, 8_000, 500
p_in, p_out, rho = 5, 25, 0.1
cost = (c * rho * p_in + (n_in - c) * p_in + n_out * p_out) / 10**6
round(cost, 4)  # → 0.0265

Two facts fall out: output tokens cost several times more than input (ask for concise answers), and a cached prefix is nearly free (put stable content first; see primer.agents.context). Caches usually charge a small premium the first time a prefix is written; this module ignores it for simplicity.

In code: request_cost is the formula above, reading each model's Price from PRICES.

Chapter 2

Route each task to the cheapest model that can do it

Everyday picture A hospital triage nurse: sprained ankles go to the nurse practitioner, chest pains to the cardiologist. Nobody sends every patient to the most expensive specialist.

Worked example The sample workload is 10 tasks: 7 simple (classify a ticket, extract a date) and 3 complex (plan a migration). All on the large model it costs $0.671. Routing the 7 simple ones to the small model, which is 5x cheaper per token, brings it to $0.314: 53% saved, with the hard tasks still on the strong model.

Figure 1 · Diagram

Reading it: the router is cheap (a few rules here; often a small classifier) and sits in front of every request. The whole saving comes from the left branch: most traffic in real systems is simple, and simple work runs well on small models. Validate the router with evals (primer.agents.evals): a router that sends hard tasks to the small model saves money and quietly loses quality.

In code: route is the rule-based router; workload_cost prices a list of WorkItem tasks with any levers switched on; routing_savings compares all-large against routed on SAMPLE_WORKLOAD.

Chapter 3

Response caching and semantic caching

Everyday picture A receptionist who's been asked "what's the wifi password?" a hundred times just answers from memory. That's an exact response cache: the same question (ignoring case and spaces) returns the stored answer with no model call. A semantic cache goes further and answers similar questions from memory, using embeddings (vectors where closeness means similar meaning, see primer.ml.embeddings) to decide "similar". That's where the receptionist starts giving the vacation-policy answer to someone asking about sick leave.

Worked example With the toy embedder in primer.common.embedder:

Stored question New question Similarity Same intent?
How do I reset my password? I forgot my password, how do I recover it? 0.88 yes
How many vacation days do I get? How many sick days do I get? 0.97 no
Request a new laptop Request a new monitor 0.96 no
What does ERR-4012 mean? What does ERR-4013 mean? 0.56 no

The two near-misses score higher than a genuine paraphrase. No single threshold separates them. Two cheap guards help: identifiers and numbers must match exactly (ERR-4012 ≠ ERR-4013), and negation must match ("cancel" ≠ "don't cancel"). They can't catch "sick" vs "vacation", which share a topic and differ in a single word.

Figure 2 · Diagram

Reading it: the exact cache is checked first because it's free and never wrong. The semantic path adds two gates: the similarity threshold, and the guards. An expired entry (older than its TTL, time to live) is treated as a miss, so answers about changing facts don't go stale forever.

Figure 3 · Drawn from the lesson's code

0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 similarity threshold 0.0 0.2 0.4 0.6 0.8 share of pairs Semantic cache: no threshold separates them cleanly paraphrases served from cache (good) different questions served from cache (wrong)

No threshold separates them: wrong hits stay at 40 to 60% through 0.95, often above the paraphrase hit rate, and reach 0 only where paraphrase hits do too

Reading it: the x-axis is the similarity threshold; the blue line is the share of genuine paraphrases answered from cache (good), and the red line is the share of different-intent questions answered from cache (a wrong answer served confidently). Lowering the threshold raises both, but look where the lines sit: from 0.4 upward the red line is above the blue one, so the cache serves more wrong answers than right ones (60% of different-intent questions against at most 57% of paraphrases). Even at 0.95 it still answers 40% of them wrongly and only 14% of paraphrases, because "sick" vs "vacation" scores 0.97. On everyday questions no threshold makes semantic caching safe; it's safe only for narrow, curated FAQ-style traffic, with guards, a TTL, and a measured wrong-hit rate.

In code: ResponseCache is the exact cache. SemanticCache is the semantic path: SemanticCache.lookup skips expired entries and entries whose identifiers or negation differ, then returns the nearest survivor, and SemanticCache.get applies the threshold. semantic_cache_sweep draws the figure from the labelled pairs in CACHE_PAIRS.

Chapter 4

Trim tokens

Everyday picture Don't photocopy the whole binder when the colleague needs one page. Compress tool results to the fields the next step needs (primer.agents.context.compress_tool_output), send the top few reranked chunks instead of dozens, remove repeated boilerplate from system prompts, and ask for concise output, because output is the expensive direction.

In code: workload_cost models trimming as cutting each task's tool output to 500 tokens, and concise output as cutting long answers by a third.

Chapter 5

Run independent tool calls at the same time

Everyday picture Boil the pasta while the sauce simmers. Cooking them one after the other takes the sum of the times; together, the time of the slowest.

Worked example Three independent lookups of 100 ms each: sequentially about 300 ms; concurrently about 100 ms.

Figure 4 · Diagram

Reading it: in the top half each call waits for the previous result; in the bottom half the three requests leave together and the agent waits only for the slowest. Models can ask for several tools in one turn (parallel tool use); run them concurrently and return all results in one message. Parallel calls cut wall-clock time, not tokens.

Figure 5 · Drawn from the lesson's code

0 50 100 150 200 250 300 milliseconds since the agent started the calls sequential parallel Three independent 100 ms tool calls tool 0 tool 1 tool 2 tool 0 tool 1 tool 2

Three 100 ms tool calls take 300 ms end to end when run one after another, but only 100 ms when started together

Reading it: each bar is one 100 ms tool call on a shared time axis. The sequential calls stack end to end, 300 ms in all; the parallel ones start together, so the whole batch finishes in the time of one. Run the lesson to measure it with asyncio: the real timings land within a few milliseconds of these bars.

In code: run_sequential awaits each call before starting the next; run_parallel starts them all with Python's asyncio gather and waits once.

Streaming (showing tokens as they're generated) doesn't reduce total time either, but users see progress immediately, which changes how fast the system feels.

Chapter 6

Batch what isn't interactive

Everyday picture Sending the laundry out to be done overnight at half price instead of waiting at the express counter. Providers offer batch APIs: submit many requests, get results within hours (often within 24), typically at about half the price. Use them for anything no one is waiting on: nightly evals, backfilling document processing, bulk classification.

In code: batch_cost prices a list of requests at the batch discount.

Chapter 7

Budgets and alerts

Everyday picture A prepaid card with a hard limit, and a bank that texts you when a purchase looks nothing like your usual spending.

A task budget caps steps and tokens per task, so a confused agent stops instead of looping all night. A tenant spend cap (a tenant is one customer organisation on a shared platform) stops one customer's runaway usage from becoming a surprise invoice. An anomaly alert fires when a task uses far more tokens than usual:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
tokens used by the task just finished 10,000
mean tokens per task over this tenant's recent history (Greek mu) 1,000
standard deviation of that history: the typical distance from the mean (Greek sigma) ≈ 71
how many standard deviations count as unusual 3

In words: alert when a task uses more tokens than the usual amount plus three times the usual spread.

On the worked example: history 1000, 1100, 900, 1050, 950, 1000 has mean 1,000 and standard deviation ≈ 71 (the squared distances from the mean add up to 25,000; divided by 6 − 1 = 5 that's 5,000, whose square root is 70.7), so the line is 1,000 + 3 × 70.7 ≈ 1,212. A 10,000-token task is far above it: alert. Such jumps usually mean a loop or a bad deploy.

Level 3: in Python
import statistics
history = [1000, 1100, 900, 1050, 950, 1000]
mu = statistics.mean(history)
# divides by 6 - 1, as above
sigma = statistics.stdev(history)
mu, round(sigma, 1)  # → (1000, 70.7)
z = 3
# the alert line
round(mu + z * sigma)  # → 1212
# alert?
10_000 > mu + z * sigma  # → True

In code: TaskBudget.charge counts each step's tokens and raises BudgetExceeded at either limit; TenantSpend.record adds a finished task's cost to its tenant's total and returns a cap alert or an anomaly alert (the formula above).

Chapter 8

Unit economics: cost per successful task

Everyday picture A cheap printer that jams on 40% of pages isn't cheap once you count the wasted paper and the time spent clearing jams.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Small model Large model
model cost of one attempt $0.002 $0.010
chance an attempt succeeds 0.60 0.95
expected number of attempts until one succeeds 1.67 1.05
cost of a person fixing one failure (illustrative: a few minutes of staff time) $2.00 $2.00

In words: if you can simply retry, you pay one attempt's cost for every expected attempt; if failures need a person, every task pays its attempt plus the chance of failure times the cost of the fix.

On the worked example: retrying, the small model costs 0.002 / 0.60 = $0.0033 per success and the large one 0.010 / 0.95 = $0.0105, so the small model wins. With human cleanup, the small model costs 0.002 + 0.40 × 2.00 = $0.802 per task and the large one 0.010 + 0.05 × 2.00 = $0.110: the "cheap" model is 7x more expensive.

Level 3: in Python
def per_success(c, p):
    # retry until it works
    return c / p
def per_task(c, p, h=2.00):
    # a person fixes each failure
    return c + (1 - p) * h
round(per_success(0.002, 0.60), 4), round(per_success(0.010, 0.95), 4)  # → (0.0033, 0.0105)
round(per_task(0.002, 0.60), 3), round(per_task(0.010, 0.95), 3)  # → (0.802, 0.11)
# small vs. large, with cleanup
round(per_task(0.002, 0.60) / per_task(0.010, 0.95))  # → 7

Figure 6 · Drawn from the lesson's code

retry until success person fixes failures 1 0 − 2 1 0 − 1 1 0 0 dollars per successful task (log scale) The cheap model is only cheap if failures are free $0.003 $0.011 $0.802 $0.110 small: $0.002/attempt, 60% success large: $0.010/attempt, 95% success

With automatic retries the small model is cheaper per success (0.3 vs 1.1 cents); if a person fixes each failure the large model wins, 11 cents vs 80

Reading it: the left pair of bars assumes failures can be retried automatically: the small model wins. The right pair assumes a person has to fix each failure: the large model wins by a mile. Which world you're in decides which model is cheaper, and the model's price per token barely matters in the second one.

In code: cost_per_success_with_retries is and cost_per_success_with_cleanup is .

Chapter 9

Putting it together: a 5x plan, in order

Order the levers so the ones that can't hurt quality come first:

  1. Prompt caching: reorder the prompt so the stable prefix is reused. No behaviour change.
  2. Trim tokens: compress tool results and send fewer chunks.
  3. Concise output: ask for shorter answers where length adds nothing.
  4. Route easy tasks to the small model: the first lever that can change quality, so it's validated with evals before and after.
  5. Batch the non-interactive share.

Figure 7 · Drawn from the lesson's code

baseline + caching + trim + concise + routing + batch 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 workload cost ($) Levers applied in order (blue: can't change quality; orange: validate with evals) 1.0x 2.2x 3.6x 3.9x 7.8x 8.4x

Workload cost falls from 67 to 8 cents as levers stack: caching 2.2x, trimming 3.6x, concise output 3.9x, routing 7.8x, batching 8.4x

Reading it: each bar is the cost of the same 10-task workload after applying every lever up to and including that one; the label is the cumulative reduction. Caching alone roughly halves it, trimming and concise output take it past 3x, routing takes it past 7x, and batching adds the last few percent. The three levers after the baseline (caching, trimming, concise output) are the free wins that don't touch quality.

In code: five_x_plan switches the levers on one at a time, in this order, and reports the cost and cumulative reduction after each.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1How would you cut the cost per task by 5x without hurting quality?Think it through, then reveal

First measure: cost per successful task, broken down by step, model, input vs. output, and cached vs. uncached tokens, with an eval set to hold quality fixed. Then, in order: restructure prompts for prompt caching; compress tool results and send only the top reranked chunks; ask for concise output; route simple steps (classification, extraction, formatting) to a small model, checking the eval before and after; move anything non-interactive to a batch API; cap steps and tokens per task. On the sample workload here that's about 8x, and every step after the first three is checked against the eval set.

Question 2Why can a cheaper model cost more?Think it through, then reveal

Because failures cost money too: retries, human cleanup, lost customers. At 60% success with a $2 human fix, a $0.002 call costs $0.80 per task; a $0.010 call at 95% costs $0.11.

Question 3What are the risks of a semantic cache?Think it through, then reveal

Returning a confident, wrong answer to a question that's similar but not the same ("sick days" vs "vacation days"), and serving stale answers after facts change. Mitigate with high thresholds, exact-match guards on IDs, numbers and negation, per-intent scoping, a TTL, and a measured wrong-hit rate on labelled pairs before enabling it.

Question 4Your average tokens per task doubled overnight. What do you check?Think it through, then reveal

Whether a deploy changed a prompt or a tool (a bigger tool output, a new retrieval setting), whether an agent is looping (the same tool called with the same arguments), and whether the prompt cache hit rate dropped (something volatile moved to the top). Traces make this a lookup rather than a guess (primer.agents.observability).

Primary sources

The papers behind this lesson

Hinton, Vinyals & Dean, Distilling the Knowledge in a Neural Network (2015), Showed how to train a small model to imitate a large one, the technique behind making the cheap model good enough to take more of the routed traffic.

Read the annotated companion →The paper ↗

Chen, Zaharia & Zou, FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance (2023), Studied prompt adaptation, caching and cascades that try cheap models first and escalate to expensive ones only when needed.

The paper ↗

Researcher's shelf

Further reading

  • Anthropic prompt caching docs: https://docs.claude.com/en/docs/build-with-claude/prompt-caching
  • Anthropic Message Batches docs: https://docs.claude.com/en/docs/build-with-claude/batch-processing
  • Anthropic, parallel tool use: https://docs.claude.com/en/docs/agents-and-tools/tool-use/implement-tool-use
  • Python asyncio.gather: https://docs.python.org/3/library/asyncio-task.html#asyncio.gather

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.