rumblr Work in progressWIP

● The AI Primer · Lesson 44 · Part 2: building systems people rely on

Retrieval-augmented generation (RAG), end to end

You'll be able to explain RAG end to end, with citations and access control

Free lesson 29 min11 figures and diagrams11 interactive
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. RAG retrieves relevant passages at question time and has the model answer from them with citations: knowledge that changes or must be cited, with no retraining.
  2. Quality is decided early: parsing (tables!), chunking with metadata, and hybrid retrieval with reranking.
  3. Filter by permissions before retrieval, never after generation.
  4. Upgrades: query rewriting, HyDE, multi-query, date filters, contextual retrieval, parent-child; agentic RAG for multi-part questions; GraphRAG for questions about connections.
  5. Debug by measuring retrieval (recall@k) first, then generation.

Level 1

The practitioner's guide

In one sentence

Retrieval-augmented generation (RAG) fetches the few passages of your own documents most likely to answer a question, puts them in the prompt, and has the model answer from those passages with citations, so the model can use knowledge it was never trained on and a reader can check where each fact came from.

When you need it

When the answer lives in documents the model has not seen (your policies, tickets, contracts, product manuals), when those documents change faster than you could retrain, when every answer must name its source, or when different users may read different documents. You don't need it when the model already knows the subject (public, stable knowledge), when you want to change how the model behaves rather than what it knows (that is fine-tuning: it teaches format and style, not facts that change), or when the whole collection fits in the prompt. Anthropic's contextual retrieval post puts that last threshold at about 200,000 tokens, roughly 500 pages: below it, put everything in the prompt and skip the pipeline. The tell: your model answers a question about the 2023 travel policy confidently and wrongly, because it never saw the 2026 one.

Your options

From the cheapest to the most capable, each one usually added on top of the last:

Option What it does What it guarantees What it costs Where it lives
Everything in the prompt Sends the whole collection with every question Nothing is ever missed by a search Every token, every call; only up to a few hundred pages Your prompt
Keyword search (BM25) Ranks passages by the question's rarer words Exact identifiers (ERR-4012, ticket numbers) are found An inverted index; misses synonyms and paraphrase A search engine
Dense search Ranks passages by embedding similarity, meaning over words Paraphrases are found ("scam message" finds the phishing guide) An embedding model at ingest and per query, a vector index; blurs exact codes A vector index
Hybrid search Runs both and fuses the rankings by rank (RRF), never by score Both kinds of question, and no score scales to reconcile Two indexes, two searches per question Both indexes plus a fusion step
Reranking A cross-encoder reads the question next to each of the top candidates Far better ordering of the top few One model score per candidate, so only over a shortlist A reranking model between search and the prompt
Retrieval upgrades Contextual chunks, query rewriting, HyDE, multi-query, date filters, parent-child Each fixes one named failure Model calls at ingest (context lines) or per query (rewrites, drafts) Ingestion or query time
Agentic RAG The model decides when, how often and with what words to search Multi-part and vague questions handled; no search for small talk Several model calls and a step budget per question An agent loop with a search tool
GraphRAG Extracts entities and relations first, answers by walking the graph or summarizing clusters Questions about connections and whole-collection themes A model pass over every document at ingest, and a graph to maintain Ingestion, plus a graph store

How to choose

Start from the questions people actually ask, with a labelled set of at least a dozen where you know the right document.

  • Questions full of exact identifiers and paraphrases both, which is what enterprise data looks like: hybrid search with a reranker. In this lesson's toy set, dense search alone never finds ERR-4012 and plateaus at 92% recall; hybrid reaches 100% by the second result; reranking the hybrid shortlist lifts the share of questions answered by the first result from 75% to 92%.
  • Chunks that lose their meaning when cut from the document (a fix that never names the error it fixes): contextual chunks. Prefixing each chunk with where it comes from turns "not in the top 20" into "second" here; Anthropic measured a 49% drop in top-20 retrieval failures from contextual embeddings and contextual BM25, and 67% with a reranker added.
  • Follow-up questions and vague phrasing: query rewriting and HyDE, which searches with a model-written draft answer.
  • Two questions in one, or a question the first search doesn't settle: agentic RAG.
  • "What depends on X?" and "what are the themes?": GraphRAG, and only then, because it is the most expensive to build and keep current.
  • Whatever you pick: filter by permissions and date before ranking, and verify citations before the answer reaches anyone. Retrieval quality is decided in the first boxes (parsing, chunking, search), never by the prompt.

What it costs

Ingestion costs an embedding per chunk, a keyword index, and for contextual chunks a model call per chunk: with prompt caching, Anthropic's post prices that at $1.02 per million document tokens. A query costs one embedding, two searches, a reranker score for each of a handful of candidates (8 here, kept to 3), and one generation whose input is the question plus those passages; that prompt is what the answer's latency and token bill mostly consist of. Agentic RAG multiplies the model calls by the number of searches (two for the two-part question in this lesson, zero for "thanks, that's all!") and needs a step limit. Quality costs come from the front of the pipeline: a table flattened by a naive PDF extractor, so that the hotel cap floats free of its grade, is a failure no later stage can repair. Effort goes mostly into the labelled question set and the parsing, not the model.

What breaks

  • The right passage is not in the index. Parsing garbled it, chunking split it from its context, or the permission sync dropped it. Check ingestion before touching anything downstream.
  • The right passage is in the index but not in the top k. Measure recall@k on the labelled set. If it is low, change retrieval (hybrid, reranker, rewriting, HyDE, contextual chunks); a prompt cannot fix it.
  • Permission leaks. Filtering after generation strips the citation and leaves the fact: in this lesson the restricted forecast's "41 million dollars" reaches an employee who may not read it. Copy each document's access list onto its chunks and filter before ranking; sync revocations promptly.
  • Stale sources. A superseded 2023 policy outranks the current one. Carry the date on every chunk and filter or prefer by it.
  • Confident answers with no support. Instruct the model to answer only from the sources and to decline otherwise, then check that every cited id was in the context and that its passage supports the claim.
  • Blaming the model. With dense-only retrieval this lesson's twelve questions produce five wrong answers, of which only one is a retrieval miss; switching to hybrid removes that one, and the two that remain are generation and data problems. Each needs a different fix, so diagnose first.

In the wild

The name comes from Lewis et al. (2020), who paired a retriever with a generator over a document index. Keyword search is BM25 in Lucene, Elasticsearch and OpenSearch; dense search runs on vector indexes such as FAISS and hosted services such as Pinecone; Elasticsearch documents RRF for fusing the two. HyDE is Gao et al. (2022). Contextual retrieval and prompt caching are described in Anthropic's post, and Claude's citations feature returns the exact passage behind each claim for PDF, plain-text and custom documents. GraphRAG is Microsoft's open implementation of Edge et al. (2024). RAGAS evaluates a RAG pipeline on retrieval and generation separately, which is exactly the split the debugging section relies on.

Go deeper

Level 2 builds the whole pipeline in plain Python: a PDF extractor's damage repaired step by step, chunks that carry their metadata, BM25 and dense search fused by a one-line formula, a reranker, a prompt that tags every source, a citation check, the permission leak reproduced, each upgrade as a working function with its before and after, and the diagnosis figure you can rerun on your own questions. If you only needed to choose, you are done.

Level 2

How it works, from scratch

What follows builds every box of the pipeline from scratch, in the order a question travels through them, then the upgrades, then how to tell a retrieval failure from a generation failure.

An open-book exam with a librarian. Before you answer, the librarian fetches the few pages most likely to contain the answer; you answer from those pages and write down which page each fact came from. If the pages don't contain the answer, you say so instead of guessing.

Retrieval-augmented generation is that arrangement for a language model. The model's training knowledge is frozen and doesn't include your company's documents, so before each answer the system retrieves relevant passages from your documents and puts them in the prompt, and the model generates an answer from them, with citations. It's how you give a model knowledge that changes, or that must be cited, without retraining it.

Figure 1 · Diagram

Reading it: the top row runs once per document, whenever documents change; the bottom row runs for every question. Everything the answer can contain has to survive every box, left to right: a table mangled at Parse or a passage missing at Retrieve can't be fixed by a better prompt at Generate. That's why most RAG quality work happens in the first boxes, not the last.

In code: RAGIndex is the top row: it chunks every document and builds the keyword and dense indexes. answer is the bottom row, from filtered retrieval to verified citations, and returns a RAGResult.

The rest of this lesson walks the boxes in order, then covers the upgrades and how to debug a wrong answer.

The librarian's shelf, working: take a document off it and ask again.

Chapter 1

Parsing: where quality dies first

Repair the extraction yourself first; then read what each repair is for.

Everyday picture Photocopy a newspaper page and read it straight across: you get the first line of one column, then the first line of the next story, and nothing makes sense. Text extracted from PDFs has the same problems: running headers and page numbers mixed into the text, words broken across lines, and tables flattened into a row of numbers with no columns.

Worked example MESSY_PDF is two pages of a travel policy as a naive extractor returns it. clean_extracted_text repairs it in four steps:

Damage Before After
Header repeated on every page Northwind Ltd Travel Policy 2026 … CONFIDENTIAL removed (it appears on 2 of 2 pages)
Page-number footer Page 1 of 2 removed
Word hyphenated across a line break all em- / ployees travelling all employees travelling
Table flattened L1-L3 150 60 L4-L6 200 75 kept as rows, then one sentence per row

The table is the dangerous one. Flattened, "200" floats free of its grade. table_rows_as_sentences turns each row into a self-contained sentence, Grade L4-L6: hotel cap 200, meal cap 75., so any row can be retrieved on its own and still says what its numbers mean.

Figure 2 · Diagram

Reading it: each box removes one kind of extraction damage. The diamond is the important decision: prose and tables need different handling, because a table's meaning lives in its columns. Real pipelines add layout-aware parsers and OCR (optical character recognition: reading text from an image of a page) for scanned documents, with a quality check on the output.

Why it matters Enterprise documents are scanned, multi-column, full of tables that span pages. If parsing garbles the numbers, no retriever or model can recover them, and the failure looks like a model error.

In code: extract_table finds a table by its header row in the cleaned text and returns each row as a dict, ready for table_rows_as_sentences.

Chapter 2

Chunking, with metadata that travels

Everyday picture Index cards. Each card holds one idea, small enough to match a question precisely, and in its corner a label: which document, which date, who may read it.

Worked example With a 15-word limit, it-004 splits into two passages:

Passage id Text Updated Readers
it-004#0 ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. 2026-04-10 everyone
it-004#1 Update AnyConnect to version 5.1 or later and reboot. 2026-04-10 everyone

(The limit is tiny because the toy documents are tiny; real systems use a few hundred tokens.) Split on the document's structure (sentences, paragraphs, sections), not at fixed character counts that cut a sentence in half. primer.ml.embeddings.retrieval compares chunking strategies in depth.

Why it matters Small chunks match precisely but can lose context (the second card doesn't say which error it fixes; see contextual retrieval below). The metadata is what makes permission and date filters possible later.

In code: chunk_document splits a document on sentence boundaries into Passages, and each Passage carries its id, title, department, date and readers.

Chapter 3

Retrieval: keyword and meaning, fused

Watch two searches disagree, and one list come out of them, before the formula.

Everyday picture Two librarians: one searches the catalogue for your exact words, the other understands what you mean even when you use different words. You take the books both of them rank highly.

  • Keyword search (BM25) scores passages by the question's words, giving more weight to rare words and less to long passages. It nails exact identifiers like ERR-4012 and misses synonyms.
  • Dense search compares embeddings (vectors where closeness means similar meaning; see primer.ml.embeddings). It finds "scam message in my inbox" → the phishing guide with no shared words, but blurs exact codes.
  • Hybrid search runs both and merges the two rankings with reciprocal rank fusion (RRF), which only needs the ranks, never the scores. BM25 scores and cosine similarities live on different scales, so adding them would be meaningless.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Worked example
one passage it-004#0
one of the rankings being fused (keyword, dense) 2 rankings
position of in ranking (1 = best); rankings that miss contribute nothing 1st, 3rd
a damping constant; 60 is the usual choice 60
add up over the rankings

In words: each ranking gives a passage a vote worth one over sixty-plus its position, and the votes are added.

On the worked example: a passage ranked 1st by dense search and 3rd by keyword search scores 1/61 + 1/63 = 0.0323. A passage ranked 1st by one list and absent from the other scores only 1/61 = 0.0164. Passages that both methods agree on rise to the top.

Level 3: in Python
k = 60
# 1st by dense, 3rd by keyword: Σ_i 1/(k + rank_i(d))
round(1 / (k + 1) + 1 / (k + 3), 4)  # → 0.0323
# 1st in one list, absent from the other
round(1 / (k + 1), 4)  # → 0.0164

Figure 3 · Drawn from the lesson's code

1 2 3 4 5 k (passages kept) 0.0 0.2 0.4 0.6 0.8 1.0 recall@k on the labelled questions Which retrieval finds the answer? keyword (BM25) dense hybrid (RRF) hybrid + rerank

Dense search plateaus at 92% recall because it never finds the error code, while hybrid reaches 100% by k = 2 and reranking lifts its top-1 from 75% to 92%

Reading it: the x-axis is how many passages you keep (k); the y-axis is recall@k, the share of the 12 labelled questions whose correct document is among those k. On this small set many questions reuse the documents' own words, so keyword search is already strong. Dense search plateaus at 92% because it never finds ERR-4012. Hybrid reaches 100% by k = 2 but only 75% at k = 1. Reranking the hybrid shortlist lifts the top-1 result to 92%. Run this measurement on your questions before choosing a method.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here
the labelled questions; is how many (12 here)
one question
1 if true, 0 if not

In words: the share of questions for which the right document made it into the top k.

On the worked example: dense search at k = 3 finds the right document for 11 of 12 questions (all but ERR-4012): 11/12 = 0.92.

Level 3: in Python
# 𝟙[...] for each of the 12 questions; ERR-4012 is the 0
found = [1] * 11 + [0]
# (1/|Q|) Σ over q in Q
round(sum(found) / len(found), 2)  # → 0.92

In code: RAGIndex.retrieve ranks the allowed passages with primer.ml.embeddings.retrieval.BM25, with dense similarity, or with both fused by primer.ml.embeddings.retrieval.reciprocal_rank_fusion, depending on the mode you ask for. recall_curve measures recall@k for each method.

Chapter 4

Reranking: read the question and each passage together

Everyday picture The librarians bring back twenty books quickly; an expert then reads your question next to each book and picks the best three.

A reranker (usually a cross-encoder: a model that reads the question and one passage together and outputs a relevance score) is far more precise than comparing precomputed vectors, and far too slow to run over every passage. So retrieval casts a wide, cheap net (here 8 candidates) and the reranker keeps the best few (here 3). This module reuses the toy cross-encoder from primer.ml.embeddings.retrieval.

In code: RAGIndex.rerank scores each (question, passage) pair with primer.ml.embeddings.retrieval.CrossEncoder and keeps the best few.

Chapter 5

Assemble, generate, cite, verify

Worked example "What does ERR-4012 mean?" becomes this context (each source escaped and tagged with its id; the question last; see primer.agents.context):

<sources>
<source id="it-004#1" title="Error ERR-4012: VPN tunnel failed" updated="2026-04-10">Update AnyConnect to version 5.1 or later and reboot.</source>
<source id="it-004#0" title="Error ERR-4012: VPN tunnel failed" updated="2026-04-10">ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date.</source>
<source id="it-002#2" title="Password policy" updated="2025-11-20">This policy applies to all employees and contractors.</source>
</sources>
<question>What does ERR-4012 mean?</question>

and the answer is "ERR-4012 means the VPN tunnel could not be established, usually because the client is out of date. [it-004#0]". verify_citations then checks that every cited id was really in the context and that the cited passage supports the claim before it; a question the sources don't cover gets "I don't know based on the provided sources."

Figure 4 · Diagram

Reading it: time runs downward. Note the order: the permission filter runs at the index, before any passage is fetched; the model only ever sees three passages; and nothing reaches the user until the citations check out. Citations are what let a person verify an answer in seconds, which is most of what makes people trust the system.

In code: assemble_context builds the tagged sources and the question. generate sends them to the model with the answer-from-sources rules, and grounded_policy is the offline stand-in model that honours them. answer chains retrieve, rerank, assemble, generate and verify_citations.

The last step above, run on an answer you can edit.

Chapter 6

Permission-aware retrieval: filter before, never after

Try to read the forecast without being in finance, both ways.

Everyday picture A librarian checks your library card before fetching from the restricted archive. Checking it after you've read the document is pointless.

Each passage carries its access-control list (ACL: the groups allowed to read it), copied from the source system (SharePoint, Google Drive, a wiki). RAGIndex.retrieve drops every passage the user can't read before ranking, so restricted text can't reach the context, the model, or the answer.

Worked example "What is the Q3 revenue forecast?" (the forecast document is readable only by finance and exec):

Who asks How it's filtered Answer
finance user before retrieval "Q3 revenue is forecast at 41 million dollars… [fin-006#0]"
anyone else before retrieval "I don't know based on the provided sources."
anyone else after generation (citation removed) "Q3 revenue is forecast at 41 million dollars…"

Figure 5 · Diagram

Reading it: in the red box the model saw the restricted passage, so the number is already in its words, and deleting the citation marker afterwards leaves the fact behind. In the green box the passage never left the index. Permission changes in the source system must also sync to the index quickly, or a revoked user keeps access until the next re-index.

In code: post_generation_filter_answer is the wrong way, kept so the leak can be shown: it retrieves for every group, generates, then strips the restricted citation markers.

Chapter 7

Retrieval upgrades

Each fix switched off and on, on the question it was made for; the table below lists them all.

Each upgrade below is a working function, and each fixes a failure you can reproduce in demo():

Upgrade Everyday picture Before → after (top 3)
Query rewriting (rewrite_query) Repeating the earlier topic when you ask a follow-up "what if it fails?" after a VPN question: printer guide → VPN setup and ERR-4012
HyDE (hyde_search) Telling the librarian "I'm looking for a page that says something like…" "a weird message asking for my bank details": nothing relevant → phishing guide first
Multi-query (multi_query_search) Asking three librarians in three different phrasings "two step login setup": no MFA guide → MFA guide in top 3
Metadata filter (updated_after) Ignoring the out-of-date binder on the shelf superseded 2023 travel policy first → gone
Contextual retrieval (RAGIndex(contextual=True)) Writing the chapter title on every index card "how do I fix ERR-4012": fix sentence missing → found
Parent-child (retrieve_parents) Find the sentence, hand over the whole page a matching sentence → its full document

HyDE (Hypothetical Document Embeddings) deserves a picture, because it sounds backwards: you ask a model to invent an answer, then search with it.

Figure 7 · Diagram

Reading it: questions and answers are written differently. The user's words ("weird", "bank details") appear nowhere in the phishing guide, but the drafted answer uses the vocabulary answers use ("phishing", "report", "suspicious"), so its embedding lands next to the real passage. The draft may be wrong in its details; that's fine, because it's only used to search and is never shown to the user.

Contextual retrieval fixes the "second index card" problem from the chunking section: before indexing, each passage gets a short line saying where it comes from. Here that's just the title and department; Anthropic's version asks a model to write one sentence situating the chunk in its document.

Figure 6 · Drawn from the lesson's code

0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 rank of the passage containing the fix (lower is better; 21 = not in top 20) how do I fix ERR-4012 ERR-4012 what should I do resolve ERR-4012 on my laptop Contextual retrieval plain chunks chunks with a context line

Without a context line the fix passage misses the top 20 for all three phrasings; with the title prefixed it ranks 2nd, 2nd and 4th

Reading it: three ways of asking how to fix ERR-4012; each pair of bars is the rank of the passage "Update AnyConnect to version 5.1 or later and reboot." (shorter is better, and 21 means not in the top 20). Without context the passage never mentions ERR-4012, and for all three phrasings it doesn't make the top 20 at all. With the title prefixed it ranks 2nd, 2nd and 4th: in the top five every time, though "resolve ERR-4012 on my laptop" still ranks three passages from the laptop guide above it. A context line turns "never found" into "found", and ranking the right passage first is then the reranker's job.

Chapter 9

GraphRAG: questions about connections

Everyday picture A detective's corkboard: photos of people and places joined by string, each string labelled with the note that established it.

Some questions aren't answered by any single passage: "what depends on two-factor authentication?" or "what are the main themes across these contracts?". GraphRAG first extracts entities (named things: systems, people, products) and the relationships between them into a graph, then answers by walking it, or by summarizing groups of connected entities. build_graph does the smallest honest version: two entities are linked when one sentence mentions both, and each link remembers which document said so.

Figure 9 · Diagram

Reading it: each circle is an entity found in the knowledge base; each line is a sentence that mentioned both ends, labelled with its source document. "What depends on two-factor?" is answered by reading the neighbours of one circle, with citations for free. Real GraphRAG uses a model to extract typed relations ("requires", "replaces") and to write a summary for each cluster, which is what makes broad, whole-collection questions answerable.

In code: communities groups the graph from build_graph into connected clusters of entities, the "themes" a real GraphRAG system would summarize.

Chapter 10

Debugging a confident wrong answer

Everyday picture A student gets an exam question wrong. Either the librarian brought the wrong book, or the student misread the right one. The fixes are completely different, so find out which before changing anything.

Worked example debug_wrong_answers runs the 12 labelled questions and classifies each: retrieval miss (no relevant document in the top 3: no prompt can fix it), generation miss (the right document was there but the answer didn't use it), or ok. With dense-only retrieval, "what does ERR-4012 mean" is a retrieval miss and recall@3 is 0.92; switching to hybrid fixes it (recall@3 = 1.00). What's left are generation misses. One of them, "per-diem for meals when traveling", is answered from the superseded 2023 policy: a data problem that a date filter fixes. The other, "enroll in MFA", is a limit of this lesson's stand-in model rather than of RAG: it matches words literally, so "enroll" misses the passage's "enrolling", and with only "MFA" in common it declines to answer. A real model would read it-006 and answer.

Figure 11 · Diagram

Reading it: start at the top and answer each question with a measurement, not a guess. Most teams jump straight to the bottom-right box and edit the prompt; the diagram says to rule out the three earlier failure points first, because they're more common and a prompt can't fix them.

Figure 10 · Drawn from the lesson's code

0 2 4 6 8 10 12 labelled questions (top 3 passages) bm25 dense hybrid Where do wrong answers come from? ok generation miss retrieval miss

Dense search has the only retrieval miss; switching to hybrid removes it, and the wrong answers left are generation misses

Reading it: each bar is one retrieval method over the 12 labelled questions, split into ok (green), generation misses (orange) and retrieval misses (red). Dense search has a retrieval miss (the error code); hybrid removes it. The orange that remains is where to look next, and it's not retrieval.

The diagnosis above, on all twelve labelled questions: change the retrieval and watch the misses move.

Test yourself

5 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Your RAG system gives confident wrong answers. Debug it step by step.Think it through, then reveal

Take the failing questions (and build a labelled set if you don't have one). First, is the right passage in the index at all? If not, it's ingestion: parsing, chunking or permission sync. Second, is it in the top k? Measure recall@k; if not, fix retrieval: hybrid search, a reranker, query rewriting or HyDE, metadata filters, contextual chunks, domain-tuned embeddings. Third, did it survive into the final context? If not, fix reranking or the context budget. Only then look at generation: instructions to answer only from sources and to decline otherwise, citation checks, stale or contradictory sources. Add every case to the golden set.

Question 2Why does hybrid search beat pure vector search on enterprise data?Think it through, then reveal

Enterprise questions are full of exact identifiers (error codes, product names, ticket numbers) that embeddings blur, and full of paraphrases that keyword search misses. Hybrid gets both, and RRF merges the rankings without having to reconcile incompatible score scales.

Question 3How do you make RAG respect document permissions?Think it through, then reveal

Copy each document's access-control list onto every chunk at ingestion, filter candidates by the user's groups before ranking, and sync permission changes from the source system promptly. Never filter after generation: the model has already used the restricted text.

Question 4When would you use agentic RAG instead of a single retrieval step?Think it through, then reveal

For multi-part or vague questions, and when the model needs to judge whether results are good enough and search again. It costs more calls and latency, so keep single-shot retrieval for simple lookups.

Question 5RAG or fine-tuning for a company knowledge assistant?Think it through, then reveal

RAG for knowledge: it updates by re-indexing, cites sources, and respects permissions. Fine-tuning teaches behaviour and format, not facts that change. Most systems are RAG plus a good prompt (see primer.ml.training_stages).

Primary sources

The papers behind this lesson

Lewis et al., Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (2020), Coined the term and showed that pairing a generator with a retriever over a document index beats a model relying on its weights alone for knowledge-heavy questions.

Read the annotated companion →The paper ↗

Gao, Ma, Lin & Callan, Precise Zero-Shot Dense Retrieval without Relevance Labels (2022), Introduced HyDE: search with an embedding of a model-written hypothetical answer.

Read the annotated companion →The paper ↗

Edge et al., From Local to Global: A Graph RAG Approach to Query-Focused Summarization (2024), Builds an entity graph and community summaries so that whole-collection questions can be answered.

The paper ↗

Researcher's shelf

Further reading

  • Anthropic, Introducing Contextual Retrieval: https://www.anthropic.com/news/contextual-retrieval
  • Anthropic citations docs: https://docs.claude.com/en/docs/build-with-claude/citations
  • Microsoft GraphRAG documentation: https://microsoft.github.io/graphrag/
  • RAGAS documentation (RAG evaluation): https://docs.ragas.io/
  • The retrieval lesson in this repo, with BM25, RRF, cross-encoders and chunking in depth: primer.ml.embeddings.retrieval

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.