rumblr Work in progressWIP

● The AI Primer · Lesson 33 · Embeddings, the centerpiece

Operating embeddings

model changes, jargon, and finding what broke

You'll be able to explain Model migrations, domain mismatch, measuring retrieval on its own

Members · open during launch 20 min7 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Every embedding model has its own vector space; never mix documents and queries from different models. Same-size models fail silently.
  2. Migrate blue/green: new index, dual-write, backfill, compare on a golden set, cut over only if better, keep the old index for rollback.
  3. Re-embedding cost is n × tokens: estimate hours and dollars before you start.
  4. General models miss company jargon; build an eval set from resolved logs, and fine-tune or add keyword search when needed.
  5. Triage RAG failures into retrieval vs. generation before fixing anything.

Level 1

The practitioner's guide

In one sentence

Operating embeddings means keeping every query on the same map as the documents it searches, changing that map without a bad day, checking that the map speaks your domain's language, and knowing which half of a retrieval-augmented system failed when an answer is wrong.

When you need it

From the day an embedding index serves real traffic. Every embedding model is its own coordinate system: in this lesson, the word "vpn" embedded by two versions of the same toy model, both 64 dimensions, has a cosine similarity of about 0.05, as if they were unrelated words. Search documents embedded by one model with queries embedded by another and recall@3 on this lesson's golden set falls from 0.92 to 0.25, barely above the 0.15 that random ranking would score, and nothing crashes. You will change models: new releases, a fine-tune, a chunking change and a bug fix all require re-embedding everything. You will meet jargon the model doesn't know: this lesson's general model puts the right document first for one of four company-jargon questions. And you will get a confident wrong answer and need to know whether retrieval or generation caused it. The tell: "we upgraded the embedding model and search got weird", or weeks spent tuning prompts for answers whose right page was never retrieved.

Your options

Two decisions recur: how to change the model, and what to do when it doesn't know your words.

Option What it does What it gives you What it costs Where it lives
Re-embed only new documents Old documents keep their old vectors; new ones get the new model Nothing: it is the classic incident. Queries and documents stop sharing a space and recall collapses silently The cheapest to run and the most expensive to discover A pipeline that forgot the backfill
Re-embed in place Stops writes, re-embeds every document into the same index, resumes A correct index at the end Downtime or a window of mixed results, and no rollback except re-embedding again Your ingestion job
Blue/green migration Builds a second index alongside, dual-writes new documents to both, backfills, compares on a golden set, then repoints a live alias No downtime, a gate that refuses a worse model, rollback by flipping one pointer Double storage during the migration, a golden set, the backfill time Your code plus index aliases (Elasticsearch, Qdrant)
Keep the general model, add keyword search Fuses BM25 with dense search so exact jargon matches by spelling Identifiers and acronyms found without training Two searches per query Any search engine (primer.ml.embeddings.retrieval)
Fine-tune on domain pairs Trains the embedding model on your own (question, document) pairs Four of four jargon questions right in this lesson, against one of four Labeled pairs, a training run, then a full re-embed primer.ml.embeddings.contrastive, followed by a migration
Measure and triage Keeps a golden set of questions with known answer passages and checks it on every change Silent failures made visible, each split into retrieval or generation A few hundred labeled questions, taken from resolved logs Your evaluation harness; Ragas

How to choose

  • Changing the model on a live index: blue/green, every time. Version the index by model, dual-write from the first minute so nothing falls behind, backfill in restartable batches, and cut over only when the new index is at least as good on your golden set. Keep the old index for a soak period.
  • Estimating the migration: multiply documents by tokens per document. Fifty million documents of 500 tokens is 25 billion tokens: about 7 hours at a million tokens per second across your workers, and $500 at $0.02 per million tokens (replace all three inputs with your own). The money is usually modest; the time, the rate limits and the double storage are what need planning.
  • The model doesn't know your jargon: measure on your own questions first, because public leaderboards measure public data. Add keyword search for exact terms, then fine-tune on domain pairs if the gap remains.
  • A wrong answer: ask one question first. Was a relevant document in the retrieved top k? If not, it is retrieval: chunking, hybrid search, reranking, the model. If yes and the answer ignored it, it is generation: the prompt, the context order, fewer distracting chunks.
  • Whatever you do, keep the golden set and rerun it on every change. It is the only number that predicts your system's quality.

What it costs

A golden set: a few hundred questions with their answer documents, taken from resolved support or search sessions, which cost nothing to collect. Re-embedding: linear in the corpus, so ten times the documents means ten times the hours and the dollars, plus double storage while both indexes exist. Blue/green adds one extra write per new document during the migration and one shadow evaluation. Fine-tuning adds labeled pairs and a training run, and then a migration, because a fine-tuned model is a new map. Triage costs a person reading a handful of failures; only the retrieval half can be measured without a model, which is why recall@k on the golden set is the first number to track.

What breaks

  • Mixed spaces. Documents from one model, queries from another, the same number of dimensions: recall falls from 0.92 to 0.25 with no error. A model with a different number of dimensions at least crashes. Tie every index to a model version and refuse vectors from any other.
  • A backfill that never finished. Only new documents got re-embedded. Dual-writing and backfilling are separate steps, and the migration is not done until every old document has been re-embedded.
  • Cutting over on faith. The shadow comparison is the gate: this lesson's migration refuses a candidate that scores 0.58 against the live index's 0.92.
  • No way back. Retire the old index only after a soak period; until then, rollback is repointing the alias.
  • Trusting a leaderboard. BEIR showed that retrieval models trained on one domain often lose badly on others, and MTEB found no single model wins everywhere. Your questions are the benchmark.
  • Tuning the prompt for a retrieval failure. The error-code question in this lesson stays wrong at every k because its page ranks fourth; no prompt fixes that. Retrieving more turns some retrieval failures into generation failures, so raising k is not a fix either.

In the wild

Index aliases are the switch: Elasticsearch's aliases API swaps an alias from one index to another in a single atomic operation with no downtime, and Qdrant's collection aliases let you build a second collection in the background and switch atomically with no concurrent request affected. Qdrant also fixes the vector size per collection, so a model with a new number of dimensions is a new collection by construction. Martin Fowler's description of blue/green deployment is where the pattern's name comes from. The MTEB leaderboard on Hugging Face ranks embedding models across tasks, and BEIR is the zero-shot retrieval benchmark that showed how much domain matters. Ragas separates retrieval metrics (context precision, context recall) from generation metrics (faithfulness, response relevancy), the same split as this lesson's triage. The papers behind this lesson are listed at the end.

Go deeper

Level 2 measures the mixed-space failure on the golden set, walks a blue/green migration state by state with the code that refuses a worse model, works the re-embedding arithmetic, shows the jargon gap and a tuned model closing it, and runs every golden question through a triage that names the failing half. If you only needed the runbook, you are done.

Level 2

How it works, from scratch

Level 2 makes each of those truths concrete, starting with two mapmakers.

The everyday picture. Two mapmakers each draw a map of the same city with their own grid. On one map, square (3, 7) is the train station; on the other, (3, 7) is a park. Both maps are fine, but you can't read a position off one map and look it up on the other.

Every embedding model is its own mapmaker. Upgrade the model and every document has to be re-plotted on the new map, and until that's done, the old map and the new one must never be mixed. That's the first of three operational truths this lesson makes concrete:

  1. Changing models means re-embedding everything, and doing it without downtime or a silent quality drop takes a plan.
  2. General models don't speak your company's jargon, so measure on your own questions and adapt when needed.
  3. When a retrieval-augmented system answers wrongly, find out which half failed: retrieval (the right page was never found) or generation (it was found and misused).

Chapter 1

Every model has its own space

Tiny worked example the word "vpn" embedded by two versions of the toy model (primer.common.embedder) with the same 64 dimensions: their cosine is about 0.06, essentially unrelated, though it's the same word. Now take 12 real questions with known answers (the "golden set" in primer.common.corpus) over 20 documents. Search v1 documents with v1 queries and the right answer is in the top 3 for 11 of 12 questions (0.92). Search the same v1 documents with queries from the new v2 model and it drops to 0.25, about what random ranking would give (3 of 20 documents shown, so 15%), and nothing crashes.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Example
Q the golden set of questions 12 questions
|Q| how many questions 12
q ∈ Q each question in turn
rel(q) the documents that truly answer q {it-004}
top_k(q) the k documents search returned 3 documents
∩, ≠ ∅ "they share at least one document"
[ … ] 1 if true, 0 if false
hit@k the share of questions with a right answer in the top k 0 to 1

In words: the share of golden questions for which at least one right document appears among the top k results. (With one relevant document per question, as here, this equals recall@k.)

On the example: 11 hits out of 12 = 0.92 with matching models; 3 out of 12 = 0.25 with mixed models.

In Python:

def hit(rel_q, top_q):
    # [rel(q) ∩ top_k(q) ≠ ∅]
    return 1 if rel_q & top_q else 0
hit({"it-004"}, {"it-002", "it-004", "fin-001"})  # → 1
# matching models: 11 of the 12 questions
hits = [1] * 11 + [0]
# (1/|Q|) Σ over q in Q
round(sum(hits) / len(hits), 2)  # → 0.92
# mixed models: 3 of the 12
hits = [1] * 3 + [0] * 9
round(sum(hits) / len(hits), 2)  # → 0.25

Figure 1 · Drawn from the lesson's code

v1 docs, v1 queries v2 docs, v2 queries v1 docs, v2 queries 0.0 0.2 0.4 0.6 0.8 1.0 recall@3 on the golden set Mixing two models' vectors fails silently 0.92 0.92 0.25 random ranking (3 of 20)

Recall@3 is 0.92 when documents and queries share a model, and falls to 0.25, near the 0.15 of random ranking, when v2 queries search v1 documents

Reading it: each bar is recall@3 on the golden set. The first two bars use one model for both documents and queries, v1 then v2: both work. The third bar searches v1 documents with v2 queries: recall falls to 0.25 (3 of the 12 questions), barely above the dashed line at 0.15, which is what random ranking would score (3 of 20 documents shown). A model with a different number of dimensions would at least crash (you can't dot a 128-number vector with a 64-number one); a same-size model fails silently, which is worse.

In code: golden_recall embeds the documents with one model and the questions with another and returns hit@k. VectorIndex is a flat index tied to one model version: VectorIndex.search embeds a text query, and VectorIndex.search_vector takes a vector that is already made.

Why it matters "we upgraded the embedding model and search got weird" is a classic incident. The cause is almost always documents and queries embedded by different models, typically because only new documents were re-embedded.

Chapter 2

Migrating without a bad day: blue/green

Everyday picture building a new bridge next to the old one. Traffic keeps using the old bridge while the new one is built and inspected. When the new bridge passes inspection, traffic is switched over, and the old bridge stays standing for a while in case something turns up.

Figure 3 · Diagram

Reading it: follow the states from the top. Users are served by v1 the whole time until cutover. Dual-writing from the very first step means documents added mid-migration land in both indexes, so the new one never falls behind. The shadow comparison is the gate: the switch only happens if v2 is at least as good on your own golden set (BlueGreenMigration refuses to cut over without it). The old index is kept after cutover, so rollback is flipping one pointer, not a multi-hour rebuild.

The "live alias" is a pointer: the search service asks for "the live index", and cutover or rollback just repoints it. Many vector databases support aliases for exactly this.

In code: each arrow of the diagram is one method: BlueGreenMigration.start, BlueGreenMigration.add (the dual write), BlueGreenMigration.backfill, BlueGreenMigration.shadow_compare, BlueGreenMigration.cutover and BlueGreenMigration.rollback.

Tiny worked example: what will re-embedding cost? 50 million documents of about 500 tokens each, an embedding throughput of 1 million tokens per second across your workers, and a price of $0.02 per million tokens (all three are inputs you replace with your own numbers):

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Example
n number of documents (or chunks) 50,000,000
t average tokens per document 500
n · t total tokens to embed 25,000,000,000
r throughput, tokens per second 1,000,000
3600 seconds per hour
p price per million tokens $0.02

In words: total tokens divided by throughput gives the time; total tokens in millions times the price gives the cost.

On the example: 25 × 10⁹ / (10⁶ × 3600) = 6.94 hours; 25,000 × $0.02 = $500.

In Python:

# documents, tokens per document
n, t = 50_000_000, 500
# tokens per second, dollars per million tokens
r, p = 1_000_000, 0.02
# total tokens
n * t  # → 25000000000
# hours
round(n * t / (r * 3600), 2)  # → 6.94
# dollars
round(n * t / 10**6 * p, 2)  # → 500.0

Figure 2 · Drawn from the lesson's code

1 0 5 1 0 6 1 0 7 1 0 8 1 0 9 documents (500 tokens each, log scale) 1 0 − 3 1 0 − 2 1 0 − 1 1 0 0 1 0 1 1 0 2 1 0 3 hours (log scale) Time to re-embed 100,000 tokens/s 1,000,000 tokens/s 10,000,000 tokens/s 1 0 5 1 0 6 1 0 7 1 0 8 1 0 9 documents (500 tokens each, log scale) 1 0 0 1 0 1 1 0 2 1 0 3 1 0 4 1 0 5 dollars (log scale) Cost to re-embed $0.01 per million tokens $0.02 per million tokens $0.13 per million tokens

Re-embedding time and cost both grow in step with corpus size: ten times the documents, ten times the hours and the dollars

Reading it: the horizontal axis is corpus size (log scale); the left panel shows hours at three throughputs, the right panel shows dollars at three prices. Both grow in straight lines on these log axes: ten times the documents, ten times the time and money. The money is usually modest; the time, the rate limits and the double storage during the migration are what need planning.

In code: reembed_estimate computes the tokens, hours and dollars for your own n, t, r and p.

Why it matters re-embedding is routine: new models, fine-tunes, chunking changes and bug fixes all require it. Versioned indexes, dual writes, a golden-set gate and a rollback path turn it from a risky event into a boring one.

Chapter 3

Domain mismatch: fluent, but not in your jargon

Everyday picture a new hire who speaks perfect English but doesn't yet know that "AP" means accounts payable or that "T&E" means travel and expenses. General embedding models are trained mostly on web text; your acronyms and product names are new words to them.

Tiny worked example four real-sounding employee questions in company jargon: "ap aging report", "hotspot from a client site", "sso lockout", "t-and-e submission deadline". The general toy model puts the right document first for 1 of 4. A version that has learned the four jargon words (DomainTunedEmbedder, standing in for a model fine-tuned on company pairs) gets 4 of 4.

Figure 4 · Drawn from the lesson's code

ap aging report hotspot from a client site sso lockout t-and-e submission deadline 0 2 4 6 8 10 12 14 rank of the right document (1 is best) Company jargon: general vs. tuned model general model tuned on the jargon

On company jargon the general model buries three of the four right answers, while the tuned model ranks every one first

Reading it: each pair of bars is one jargon question; the height is where the right document ranked (1 is best, shorter is better). The general model buries three of the four answers; the tuned model puts every one first. The one the general model gets right, "sso lockout", is saved by a word it does know ("lockout").

In code: rank_of_first_relevant finds where the right document lands for one question under one model: the height of each bar.

Figure 5 · Diagram

Reading it: the evaluation set comes from real usage, not invention. Resolved sessions give you (question, answer document) pairs for free; build_eval_set keeps only resolved ones and merges repeats. Test candidate models on that set, and when none is good enough, fine-tune on domain pairs (primer.ml.embeddings.contrastive) and add keyword search, which matches jargon exactly (primer.ml.embeddings.retrieval).

Why it matters public leaderboards measure public data. The only number that predicts your system's quality is recall on your own questions.

Chapter 4

Measure retrieval separately from generation

Everyday picture an open-book exam. A wrong answer has one of two causes: the right page wasn't in the book you brought (retrieval failure), or it was, and you misread it (generation failure). Studying harder fixes the second, not the first.

Tiny worked example three answered questions, checked by hand.

Retrieved (top 3) Truly relevant Answer used Verdict
it-002, fin-006, fin-001 it-004 it-002 retrieval failure: it-004 was never found
hr-002, hr-001, hr-005 hr-001 hr-002 generation failure: found but not used
it-001, it-002 it-001 it-001 correct

Figure 7 · Diagram

Reading it: one question splits every failure in two. If the right document wasn't retrieved, no prompt change can help, so work on retrieval. If it was retrieved and ignored, the problem is downstream, in the prompt or the context. Only the retrieval half can be measured without a model, which is why recall@k on a golden set is the first number to track.

Figure 6 · Drawn from the lesson's code

1 3 5 documents retrieved (k) 0 2 4 6 8 10 12 golden questions Which half failed? (generator uses the top document) correct generation failure retrieval failure

Retrieving more documents turns some retrieval failures into generation failures, and the error-code question stays wrong at every k

Reading it: each bar is the 12 golden questions, answered by a toy generator that always uses the top document, with k documents retrieved. Green is correct, red is a retrieval failure, orange is a generation failure. Retrieving more (larger k) turns some retrieval failures into generation failures: the right page is now in the book, but the reader still opened the wrong one. Watch the error-code question ("what does ERR-4012 mean"): its answer ranks only fourth, so it's a retrieval failure until k = 5, and even then the top document is the wrong one. Dense vectors blur exact identifiers; keyword search would put that page first.

In code: triage gives the verdict for one answered question, following the flowchart above; triage_golden_set runs every golden question through retrieval and a toy generator that answers from the top document.

Why it matters teams burn weeks tuning prompts for failures that were retrieval all along. Triage first, then fix the half that's broken.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: You need to switch embedding models on a 50-million-document index. Plan the migration.Think it through, then reveal

Create a versioned index for the new model and dual-write new documents to both. Backfill by re-embedding existing documents in restartable batches (estimate: 50M × 500 tokens = 25B tokens; at 1M tokens/s that's about 7 hours). Compare recall on a golden set in shadow; cut over the live alias only if the new model is at least as good; keep the old index for fast rollback, then retire it.

Question 2Q: Why can't you compare vectors from two different embedding models?Think it through, then reveal

Each model defines its own coordinate system. Even with the same number of dimensions, the same text lands in unrelated places, so similarity between the two spaces is meaningless.

Question 3Q: A general embedding model performs poorly on a company's internal documents. What helps?Think it through, then reveal

Build a small eval set from real queries and the documents that resolved them, test several models on it, add hybrid (keyword + dense) search for exact terms, and fine-tune an embedding model on domain pairs with hard negatives if needed.

Question 4Q: Your RAG system gives confident wrong answers. What's your first diagnostic step?Think it through, then reveal

Split retrieval from generation: for failing questions, check whether a relevant document was in the retrieved top k. If not, it's retrieval (fix search); if yes, it's generation (fix prompt and context). Track recall@k on a labeled set continuously.

Primary sources

The papers behind this lesson

Thakur et al., BEIR: A Heterogeneous Benchmark for Zero-shot Evaluation of Information Retrieval Models (2021)

Showed that retrieval models trained on one domain often lose badly on others, which is why you evaluate on your own data.

The paper ↗
Muennighoff et al., MTEB: Massive Text Embedding Benchmark (2022)

Compared embedding models across many tasks and found no single model wins everywhere.

The paper ↗

Researcher's shelf

Further reading

  • Es et al., RAGAS: Automated Evaluation of Retrieval Augmented Generation (2023): https://arxiv.org/abs/2309.15217
  • Martin Fowler, BlueGreenDeployment: https://martinfowler.com/bliki/BlueGreenDeployment.html
  • MTEB leaderboard: https://huggingface.co/spaces/mteb/leaderboard

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.