rumblr Work in progressWIP

● The AI Primer · Lesson 28 · Embeddings, the centerpiece

Contrastive training

how an embedding model learns what "similar" means

You'll be able to explain Contrastive learning, hard negatives, CLIP

Members · open during launch 26 min10 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Contrastive training pulls matching pairs together and pushes non-matches apart; the loss (InfoNCE) is softmax cross-entropy over the batch.
  2. In-batch negatives are free: every other answer in the batch.
  3. Negatives only teach the distinctions they contain. Hard negatives (same topic, wrong answer) are what teach retrieval rather than topic matching.
  4. Temperature sharpens the softmax so small cosine gaps count; typical τ is 0.01 to 0.1.
  5. CLIP applies the same loss in both directions between an image encoder and a text encoder, which gives one shared space and zero-shot classification.

Level 1

The practitioner's guide

In one sentence

Contrastive training teaches an embedding model what "similar" means by showing it pairs that belong together and pairs that don't, pulling the first kind close and pushing the second apart, and it is how nearly every embedding model behind search and RAG was made.

When you need it

You need to understand it whenever you pick or judge an embedding model, because what a model calls similar is exactly what its training pairs called similar: a model trained on (question, answer) pairs ranks answers, one trained on (sentence, paraphrase) pairs ranks restatements, and neither is "the" similarity. You need to run it when a general model keeps confusing things your users never confuse: the reset steps with the password policy, two product lines that share a name, a statute with its commentary. The tell is a retrieval log where the top hit is on topic and still wrong. This lesson's experiment puts a number on it: a model trained only with random negatives picks the right card over its look-alike 60% of the time on unseen topics, barely above a coin flip, scoring the reset card at 0.646 and the policy card at 0.632 for "my password is broken"; the same model trained with hard negatives picks right 100% of the time and scores them 0.944 and −0.588. You don't need to train anything when a general model already ranks your held-out queries well, and you don't need it for one-off comparisons where a person reads the result.

Your options

From the cheapest to the most work:

Option What it does What it guarantees What it costs Where it lives
A general pretrained model, as is Uses vectors from a model trained on millions of public pairs Good general similarity; on public benchmarks, dense models beat keyword search (DPR: 9 to 19 points of top-20 accuracy over BM25) Per-token fees or hosting, and nothing about your jargon An embedding API or an open model
The model's query and document modes Embeds each side of a search the way the model was trained for it (a query prefix, an input type) The pairing the model learned, instead of a symmetric comparison it wasn't trained for Reading the model card Your embedding calls
Fine-tune with in-batch negatives Trains on your (query, right document) pairs; every other document in the batch is a free negative Learns your vocabulary and topics from a few thousand pairs Pairs from logs, a training run, a full re-embedding of the corpus Your training job
Fine-tune with mined hard negatives Adds to each pair a top search result that is wrong Learns the distinction retrieval needs: right answer versus look-alike A mining pass over the corpus per training query, plus the above Your training job
Contrastive pretraining at scale Trains from weak public pairs before any fine-tuning (E5's CCPairs) A general model that beats BM25 zero-shot Curated web-scale pairs and GPU weeks; for model builders A research lab or a model vendor
Two encoders, one space (CLIP) Trains an image encoder and a text encoder so matching pairs meet Search across modalities and zero-shot labelling; 26% to 92% on the lesson's toy Paired data across modalities (CLIP used 400 million pairs) A multimodal model

How to choose

Start from what the model must tell apart, and whether it already can.

  • Measure first: embed a few hundred held-out queries with the general model and check recall@k against the documents that resolved them. Good enough means stop.
  • Check the model card for a query mode before blaming the model; embedding both sides the same way when it expects a prefix costs quality for free.
  • Failures are about vocabulary (your terms, your product names): fine-tune on pairs from logs with in-batch negatives.
  • Failures are about intent (right topic, wrong document): mine hard negatives from your current search results and train with them. This is the step that moved the lesson's model from 60% to 100%.
  • Images, scans or audio next to text: a CLIP-style dual-encoder model, not a text model with captions.
  • Whatever you pick, judge it on held-out pairs against the base model, and ship only a winner. Fine-tuning changes the whole space, so every stored vector is re-embedded.

What it costs

Data is the main price, and it is cheap when you have logs: real questions and the document that resolved each, support tickets and the article that closed them. A few thousand good pairs usually adapt a general model to a domain. In-batch negatives are free, which is why batches matter: a batch of B pairs gives every question B − 1 negatives at no extra cost, and sentence-transformers' guidance for its MultipleNegativesRankingLoss is that larger batches are better, with a cached variant (GradCache) for large batches on limited memory. Hard negatives cost one search per training question, with BM25 or the model itself. Compute for fine-tuning is small next to pretraining: the lesson's bi-encoder trains for 400 steps at batch 8 in seconds, and a fine-tune starts from a model whose web-scale training someone else paid for. The recurring bill is deployment: a new model means re-embedding the whole corpus, and every model change repeats it. One setting to get right: the temperature, the dial that turns small cosine gaps into confident scores; typical values are 0.01 to 0.1 (sentence-transformers' default scale of 20 is a temperature of 0.05).

What breaks

  • Topic matching instead of answering. Random negatives are about other subjects, so the model learns subjects: 60% on look-alikes. Mine hard negatives.
  • A hidden right answer in the batch. Every card that isn't a question's own is treated as wrong, so two questions with the same answer in one batch push a right answer away. Deduplicate pairs before batching; sentence-transformers' GISTEmbedLoss guides in-batch sampling for this.
  • Temperature at the wrong setting. At a temperature of 1, a right card leading three wrong ones by 0.2 in cosine earns only 29% of the softmax, so training keeps punishing correct rankings; at 0.05 it earns 95%. Too low and a few hard pairs dominate and training turns unstable.
  • Training and testing on the same topics. The lesson's exam uses four topics the model never saw; a model scored on its training topics looks better than it is.
  • The old index after a new model. A fine-tuned model is a new space. Vectors from the old one are not comparable; re-embed everything.
  • Both sides embedded the same way. A model trained with a query side and a document side scores worse when you skip the prefix or the input type it was trained with.

In the wild

Sentence-BERT (2019) made the bi-encoder the standard shape; DPR (2020) trained one on question-passage pairs with in-batch and BM25-mined hard negatives and beat BM25 by 9 to 19 points of top-20 accuracy; E5 (2022) pretrained contrastively on weakly supervised web pairs and was the first to beat BM25 on BEIR with no labels; SimCSE (2021) showed that dropout noise alone gives usable positives, with NLI pairs as hard negatives; CLIP (2021) trained an image and a text encoder on 400 million pairs with the symmetric loss. sentence-transformers implements the loss as MultipleNegativesRankingLoss, with CachedMultipleNegativesRankingLoss for big batches and GISTEmbedLoss for guided negatives. Cohere's embedding API asks which side of a search each text is on. The MTEB leaderboard ranks embedding models by the tasks these losses target.

Go deeper

Level 2 scores one question against two answer cards by hand, turns the scores into a loss (InfoNCE) and shows what the temperature does to it, then runs the controlled experiment: two identical models, one with hard negatives, tested on topics neither has seen. It ends with the CLIP recipe, two encoders trained into one space, and a zero-shot classifier that needs no classifier. If you only needed to choose, you are done.

Level 2

How it works, from scratch

A teacher has a stack of question cards and a stack of answer cards. She lays out a few questions, deals out all the answers face up, and asks the student to pair each question with its answer. Every wrong pairing, she corrects. Over many rounds the student learns what makes an answer fit a question.

Then she gets sneaky. Alongside each right answer she slips in a look-alike wrong card: same subject, wrong answer. "How do I reset my password?" now faces both the reset steps and the password policy. A student who only ever saw easy wrong cards (answers about printers or holidays) would happily pick the policy card because it says "password". The look-alikes force the student to learn the difference that actually matters.

That's contrastive training, and it's how nearly every modern embedding model is taught. The student is the model; "pairing" means making a question's vector point the same way as its answer's vector (see primer.ml.embeddings.similarity); the wrong cards are negatives; the look-alikes are hard negatives.

Chapter 1

A tiny worked example: one question, two answer cards

A question's vector is q = (1, 0). The right answer is p₊ = (1, 0) and a wrong one is p₋ = (0, 1). All three have length 1, so a dot product is a cosine.

  1. Score each card with the dot product: q·p₊ = 1, q·p₋ = 0.
  2. Divide by the temperature τ (tau), a sharpness dial explained below. With τ = 1 the scores stay (1, 0).
  3. Softmax (from primer.ml.attention): exponentiate and share out. e¹ = 2.718 and e⁰ = 1, so the right card gets 2.718 / 3.718 = 0.731.
  4. Loss = −ln 0.731 = 0.313. It would be 0 if the right card got 100%.

At τ = 0.1 the scores become (10, 0), the right card gets 99.995%, and the loss is 0.0000454. Same vectors, far more confident: that's what the temperature does.

Figure 1 · Diagram

Reading it: three texts go through the same encoder (the model that turns text into a vector). The question and its right answer are pulled together; the question and the look-alike are pushed apart. The look-alike is about the same topic but doesn't answer the question. Learning to separate it from the right answer is what makes a model good at retrieval (finding the answer) rather than just topic matching (finding the subject).

Chapter 2

The math: InfoNCE, the loss behind it

In practice the teacher deals a whole batch. B questions sit in rows, all the answer cards in the batch sit in columns, and every card that isn't a question's own answer is a free negative for it: in-batch negatives.

Figure 2 · Diagram

Reading it: each row is one question scored against every answer card in the batch. The ✔ on the diagonal is its own answer; every ✗ is a negative that costs nothing extra because those answers are already in the batch. Extra columns to the right of the bar (|) are hard negatives added on purpose, which belong to no question. Softmax runs along each row, and the loss asks each row to put its share on the ✔.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
B number of questions in the batch 8 in this lesson
M number of answer cards in the batch (B, plus any hard negatives) 8 or 16 here
i which question (row) 1 to B
j which answer card (column) 1 to M
qᵢ the i-th question's vector, unit length d numbers (d = 16)
pᵢ question i's own answer vector (the ✔) d numbers
pⱼ the j-th answer card in the batch d numbers
· dot product; equals cosine because the vectors are unit length −1 to 1
τ (tau) temperature: divides every score; small τ sharpens the softmax 0.01 to 1; 0.1 here
exp e raised to the power > 0
Σⱼ add up over all M answer cards
the fraction softmax share of the right card 0 to 1
log, −(1/B) Σᵢ take −log of each row's share, then average over the rows L ≥ 0

In words: for every question, measure what share of its softmax goes to its own answer, take the negative log (big when the share is small), and average over the batch.

On the example: B = 1, M = 2, τ = 1: L = −log(e¹ / (e¹ + e⁰)) = −log 0.731 = 0.313.

In Python:

import math
def dot(a, b):
    return sum(a_i * b_i for a_i, b_i in zip(a, b))
def info_nce(q, p, tau):
    B = len(q)
    total = 0
    for i in range(B):
        # exp(q_i · p_j / τ) for every card j
        exps = [math.exp(dot(q[i], p_j) / tau) for p_j in p]
        # log of the ✔ card's share
        total += math.log(exps[i] / sum(exps))
    # -(1/B) Σ_i
    return -total / B
# B = 1 question
q = [(1, 0)]
# M = 2 cards; card 0 is the question's own answer
p = [(1, 0), (0, 1)]
round(info_nce(q, p, tau=1), 3)  # → 0.313
print(f"{info_nce(q, p, tau=0.1):.7f}")  # → 0.0000454

The name InfoNCE comes from "noise-contrastive estimation": telling the true pair apart from noise. It's the softmax cross-entropy loss from primer.ml.losses, where the "classes" are the answer cards in the batch.

The model trained here is a deliberately tiny bi-encoder (the same encoder applied separately to questions and answers): count the words in a text (a bag of words), multiply by one learned matrix W, and scale the result to length 1. _contrastive_grads and _normalize_backward derive the gradient (the direction in which each number in W should move to lower the loss) by hand, and encoder_gradient_check confirms it against a brute-force numerical estimate.

In code: info_nce_loss computes L for a batch, treating any rows of P beyond B as extra hard negatives. BagOfWords turns text into word counts, BiEncoder holds the learned matrix W, and BiEncoder.encode maps texts to unit vectors.

Why it matters the embedding models behind search and RAG (Sentence-BERT, DPR, E5 and their successors) are trained with exactly this loss on millions of (question, answer) pairs. When a retrieval system confuses "reset my password" with "password policy", this is the training signal that was missing.

Chapter 3

Temperature: the sharpness dial

Everyday picture grading on a curve. With a gentle curve (high τ), a slightly better answer gets slightly more credit. With a steep curve (low τ), the best answer takes nearly all the credit, so small differences in score become big differences in outcome.

Tiny example the right card has cosine 0.8 and three wrong cards have 0.6 each.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Range
share₊ the softmax share that goes to the right card 0 to 1
0.8, 0.6 cosine of the right card and of each wrong card −1 to 1
3 the number of (identical) wrong cards
τ (tau) temperature: every cosine is divided by it 0.01 to 1 in practice
e^x e ≈ 2.718 raised to the power x > 0

In words: the right card's share is its exponentiated, temperature-scaled score divided by the total over all cards.

On the example: at τ = 1, 2.2255 / (2.2255 + 3 × 1.8221) = 0.289; at τ = 0.05 the scores become 16 and 12, and 1 / (1 + 3e⁻⁴) = 0.948.

In Python:

import math
def share_plus(tau):
    # e^(0.8/τ)
    right = math.exp(0.8 / tau)
    # three wrong cards, e^(0.6/τ) each
    wrong = 3 * math.exp(0.6 / tau)
    return right / (right + wrong)
round(share_plus(1), 3), round(share_plus(0.05), 3)  # → (0.289, 0.948)

Figure 3 · Drawn from the lesson's code

1 0 − 2 1 0 − 1 1 0 0 temperature τ (log scale) 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 softmax share of the right card Right card at cos 0.8 vs. three wrong at 0.6

With the right card 0.2 ahead in cosine it gets under a third of the softmax share at temperature 1 and about 95% at temperature 0.05

Reading it: the horizontal axis is τ (log scale), the vertical axis is the right card's softmax share when it beats three wrong cards by 0.2 in cosine. At τ = 1 a 0.2 lead earns under a third of the share, so the loss keeps complaining about pairs that are already ranked correctly. Around τ = 0.05 the same lead earns about 95%. Cosines live in a narrow range (−1 to 1), so embedding models use small temperatures (commonly 0.01 to 0.1) to make those small differences count.

In code: positive_probability computes share₊ for given cosines and τ; this curve is that function swept across τ.

Why it matters too high a temperature and the model can't become confident; too low and a few hard pairs dominate training and it becomes unstable. It's one of the handful of settings that really matters when fine-tuning an embedding model.

Chapter 4

Hard negatives: the experiment

Everyday picture a driving test on an empty road teaches you to steer. It teaches nothing about merging, because merging never came up. Negatives only teach the distinctions they contain.

The setup (train_bi_encoder): questions come in two intents for each topic: "how do I fix it?" (answered by a how-to card: "restart, reinstall and follow the setup guide") and "what are the rules?" (answered by a policy card: "approval, compliance and usage limits"). The question and answer cards share almost no words, so matching intent must be learned ("fix" goes with "restart"), not read off the overlap. Each batch holds one intent across different topics, so its in-batch negatives differ from the right card only by topic. We train two identical models; the only difference is that one also gets each question's same-topic, other-intent card as a hard negative. Then we test on four topics neither model has seen, including "password".

Figure 7 · Diagram

Reading it: both models see exactly the same questions and answers in the same order. The branch on the right adds one extra wrong card per question, a card that has the right topic but the wrong intent. Both models then take the same exam: for a question about an unseen topic, does the right card beat its look-alike?

Figure 4 · Drawn from the lesson's code

0 50 100 150 200 250 300 350 400 training steps 0.0 0.2 0.4 0.6 0.8 1.0 right card beats look-alike (unseen topics) Negatives only teach the distinctions they contain in-batch negatives only plus hard negatives coin flip

Without hard negatives the model stays mostly below a coin flip on look-alike cards and ends near 60%, while with hard negatives it reaches 100% within ten steps

Reading it: the horizontal axis is training steps, the vertical axis is the share of held-out questions whose right card beats the same-topic look-alike (the dashed line is a coin flip). The in-batch-only model spends most of its training below the coin flip, often preferring the wrong card, and ends near 60%. Nothing in its training ever penalised "same topic, wrong intent", so whatever it does on that distinction is an accident of what it learned about topics. With hard negatives the model reaches 100% within ten steps, on topics it never trained on, because it learned what "fix" and "rules" mean.

Figure 5 · Drawn from the lesson's code

−0.6 −0.4 −0.2 0.0 0.2 0.4 0.6 principal component 1 −0.6 −0.4 −0.2 0.0 0.2 0.4 0.6 0.8 principal component 2 in-batch negatives only −0.75 −0.50 −0.25 0.00 0.25 0.50 0.75 1.00 principal component 1 −0.6 −0.4 −0.2 0.0 0.2 0.4 principal component 2 plus hard negatives how-to question policy question how-to card policy card

With in-batch negatives only, held-out questions and cards group by topic; with hard negatives they split by intent, each question beside the card that answers it

Reading it: each panel squashes the embeddings of the four unseen topics to 2-D with PCA (the two most spread-out directions; see primer.notation). Circles are questions and squares are answer cards; blue means how-to and orange means policy. With in-batch negatives only (left), the points group by topic, and how-to and policy questions sit on top of each other. With hard negatives (right), they split by intent, and each question sits next to the kind of card that actually answers it.

Figure 6 · Drawn from the lesson's code

reset card policy card how do i fix my password my password is broken password not working help i am stuck with my password password keeps failing what are the password rules is there a limit on password who can get password password rules for my team am i permitted to have password in-batch negatives only reset card policy card plus hard negatives −0.2 0.0 0.2 0.4 0.6 0.8 1.0 cosine similarity

On the unseen password topic, the in-batch-only model scores both cards almost alike, while the hard-negative model lights up the reset card for how-to questions and the policy card for policy questions

Reading it: rows are the ten "password" questions (five how-to, then five policy), columns are the two password cards, and brighter means a higher cosine. On the left the two columns are nearly the same colour: the model sees "password" and can't tell the cards apart. On the right, the how-to rows light up the reset card and the policy rows light up the policy card: the diagonal blocks are exactly the right answers.

In code: make_pairs builds the (question, right card, look-alike) triples; look_alike_accuracy runs the exam on unseen topics, and training_curve records that score at checkpoints during training.

Why it matters hard negatives are the biggest single driver of retrieval quality. In practice you mine them: search your corpus with BM25 or the current model, and take high-ranking results that are not the right answer.

Chapter 5

Fine-tuning on your own domain

Figure 8 · Diagram

Reading it: start from pairs you already have: search logs, support tickets and the article that closed them, questions and their FAQ entries. Mine hard negatives from your current search results. Train, then evaluate recall@k (the share of questions whose right document is in the top k) on held-out pairs, and only ship if it beats the base model. Shipping means re-embedding the whole corpus, because a fine-tuned model has a new vector space (primer.ml.embeddings.operations).

A few thousand good pairs is usually enough to adapt a general model to legal, medical or internal jargon. Libraries such as sentence-transformers implement this loss as MultipleNegativesRankingLoss.

Chapter 6

CLIP: two encoders, one shared space

Everyday picture two people describing the same holiday, one through photos and one through postcards. CLIP teaches a photo reader and a text reader to put matching photos and captions at the same spot on one shared map, so a photo of a dog and the words "a photo of a dog" land together.

Tiny example two images and two captions, each a unit vector. Correctly paired, image 1 scores (1, 0) against the captions and caption 1 scores (1, 0) against the images: each direction's loss is ln(1 + 1/e) = 0.313. Swap the captions and every right pair scores 0 against a wrong pair's 1: the loss rises to ln(1 + e) = 1.313.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here
L_image→text InfoNCE where each image (a row of the table) must pick its own caption from all captions in the batch
L_text→image the same loss with the roles swapped: each caption (a column) must pick its own image
½( … + … ) the average of the two directions
L_CLIP the loss both encoders are trained on together

In words: each image must find its caption and each caption must find its image.

On the example: ½(0.313 + 0.313) = 0.313 correctly paired; ½(1.313 + 1.313) = 1.313 with the captions swapped.

In Python:

import math
# InfoNCE with τ = 1: row i must pick column i
def rows_pick_diagonal(S):
    return -sum(math.log(math.exp(row[i]) / sum(math.exp(s) for s in row))
                for i, row in enumerate(S)) / len(S)
def clip_loss(S):
    # captions picking images
    columns = [list(col) for col in zip(*S)]
    # ½(L_image→text + L_text→image)
    return (rows_pick_diagonal(S) + rows_pick_diagonal(columns)) / 2
# correctly paired
round(clip_loss([[1, 0], [0, 1]]), 3)  # → 0.313
# captions swapped
round(clip_loss([[0, 1], [1, 0]]), 3)  # → 1.313

In code: clip_loss computes L_CLIP by averaging info_nce_loss over the rows (images pick captions) and the columns (captions pick images).

Figure 10 · Diagram

Reading it: two different encoders, one per modality, each end in vectors of the same length. The similarity table compares every image with every caption in the batch, and the loss runs along the rows and down the columns. The extra input at the bottom left is the zero-shot trick: to classify an image with no extra training, write a caption for each label, embed it, and pick the label whose caption is closest.

train_clip_toy trains two small linear encoders on toy "images" and "captions" that describe the same hidden meaning through different, unrelated feature spaces. Before training, zero-shot labeling is near chance (20% for five classes); after training it's above 90%.

Figure 9 · Drawn from the lesson's code

'a photo of class 0' 'a photo of class 1' 'a photo of class 2' 'a photo of class 3' 'a photo of class 4' 0 5 10 15 20 25 30 35 test images, sorted by true class Zero-shot accuracy 92% −0.6 −0.4 −0.2 0.0 0.2 0.4 0.6 0.8 cosine similarity

After training, 36 of the 40 test images are most similar to their own label's caption, forming a diagonal staircase; the 4 misses are mostly classes 0 and 4 mistaken for each other

Reading it: rows are 40 test images sorted by their true class, columns are the five label captions, and brighter means more similar. Each block of rows lights up brightest in its own label's column, and that diagonal staircase is zero-shot classification working: no classifier was trained, just two encoders that agree on where meanings live. It isn't perfect. Look for the rows whose brightest cell sits in the wrong column: 4 of these 40 images are closest to another class's caption, three of them classes 0 and 4 mistaken for each other. Over all 500 held-out test images the accuracy is 92%, the number in the figure's title.

Why it matters the shared space enables search across modalities (find images with text, or text with images), zero-shot classification, and mixing scanned document images with text in one retrieval index. Multimodal models use the same recipe.

Test yourself

5 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: How are embedding models trained?Think it through, then reveal

Contrastively, on (query, relevant passage) pairs: an encoder embeds both, and InfoNCE rewards the pair's similarity over the similarity to other passages in the batch (in-batch negatives) and to deliberately chosen hard negatives.

Question 2Q: Why do hard negatives matter so much?Think it through, then reveal

Random negatives are usually about something else, so a model can beat them by topic matching alone. Hard negatives share the topic but don't answer the question, forcing the model to learn the fine distinction that retrieval actually needs.

Question 3Q: What does the temperature do in InfoNCE?Think it through, then reveal

It divides the similarities before softmax. A small temperature makes the softmax sharp, so small cosine differences produce large differences in probability and gradient. Too small makes training unstable; too large makes the model unable to become confident.

Question 4Q: How does CLIP put images and text in the same space?Think it through, then reveal

It trains an image encoder and a text encoder together on image-caption pairs with a symmetric contrastive loss: each image must pick its caption from the batch and each caption must pick its image. Matching pairs end up close in one shared space.

Question 5Q: How would you adapt an embedding model to a company's jargon?Think it through, then reveal

Collect real (query, correct document) pairs from logs, mine hard negatives from current search results, fine-tune with in-batch plus hard negatives, evaluate recall@k on a held-out set against the base model, then re-embed the corpus with the new model.

Primary sources

The papers behind this lesson

van den Oord, Li and Vinyals, Representation Learning with Contrastive Predictive Coding (2018)

Named and popularised the InfoNCE loss: pick the true sample out of a set of negatives.

The paper ↗
Reimers and Gurevych, Sentence-BERT (2019)

Turned BERT into a bi-encoder that produces one comparable vector per sentence, making semantic search with transformers fast.

Read the annotated companion →The paper ↗
Karpukhin et al., Dense Passage Retrieval for Open-Domain Question Answering (2020)

Showed a question/passage bi-encoder trained with in-batch negatives plus BM25-mined hard negatives beats keyword search for open-domain QA.

Read the annotated companion →The paper ↗
Radford et al., Learning Transferable Visual Models From Natural Language Supervision (CLIP, 2021)

Trained image and text encoders with a symmetric contrastive loss on 400 million pairs, giving zero-shot image classification.

Read the annotated companion →The paper ↗
Gao, Yao and Chen, SimCSE (2021)

Showed contrastive learning of sentence embeddings works even with dropout noise as the only "augmentation", and with NLI pairs as hard negatives.

The paper ↗

Researcher's shelf

Further reading

  • sentence-transformers, training overview: https://www.sbert.net/docs/sentence_transformer/training_overview.html
  • Wang et al., Text Embeddings by Weakly-Supervised Contrastive Pre-training (E5, 2022): https://arxiv.org/abs/2212.03533
  • Muennighoff et al., MTEB: Massive Text Embedding Benchmark (2022): https://arxiv.org/abs/2210.07316
  • MTEB leaderboard: https://huggingface.co/spaces/mteb/leaderboard
  • Lilian Weng, Contrastive Representation Learning: https://lilianweng.github.io/posts/2021-05-31-contrastive/

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.