The lesson in one minute
What you'll be able to explain
- Cosine compares direction, the dot product compares direction and length, Euclidean distance measures the gap between the tips.
- On L2-normalized vectors all three give the same ranking, because ‖a − b‖² = 2 − 2cos.
- Use the metric the model was trained with.
- Scores aren't comparable across models (anisotropy). Rank when you can; calibrate thresholds on labeled pairs when you can't.
Level 1
The practitioner's guide
In one sentence
A similarity metric turns two embedding vectors into one number that says how alike their meanings are, and the three in common use (cosine, dot product, Euclidean distance) agree with each other only under conditions you have to arrange.
When you need it
Every time you compare embeddings: semantic search,
retrieval for RAG, near-duplicate detection, a semantic cache that answers a
question from a stored one, clustering, recommendations. You choose a metric
whether you mean to or not, because every vector database and every search
library has a default, and the wrong one ranks silently wrong. The tells:
you are about to write ORDER BY against a vector column and have not
checked which operator the embedding model expects; or your team argues
about whether "0.8 similar" is a lot. This lesson's demo shows why the
second argument can't be settled by intuition: for one model, paraphrases
score between 0.76 and 0.87 and unrelated pairs between 0.72 and 0.78, so
"0.8" is either a near-duplicate or a stranger. You don't need any of this
for exact-match lookups (an id, a hash) or for keyword search, which scores
words, not vectors.
Your options
The ways to score a pair, from the cheapest to the most deliberate:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Dot product on normalized vectors | Normalize every vector to length 1 once, at write time, then multiply and add | The same ranking as cosine and as Euclidean distance, at the lowest price | One multiply-add per dimension per pair; a normalization step on insert | Your ingestion code, then the database's inner-product operator |
| Cosine similarity | Dot product divided by both lengths | Direction only; lengths can't skew the order, whatever came in | Two extra norms per comparison, unless the store caches them | The database's cosine operator, or a library's default |
| Raw dot product | Direction and length together | Keeps signal a model put in the length, such as popularity | Rankings that a mix of long and short vectors can dominate | Recommenders and models trained with it |
| Euclidean distance | Straight-line gap between the tips | The natural metric for clustering and for many index structures | Disagrees with cosine unless vectors are normalized | k-means, some indexes |
| Mean-centred scores | Subtract the collection's average vector before comparing | Unrelated pairs land near 0, similar ones stand clear | A pass over the collection, and a re-run when it changes | Your code, before scoring |
| A calibrated threshold | Label a few hundred pairs, sweep the cut-off, keep the best F1 | A yes/no that matches your data and your model | A labelling session, repeated at every model change | A number stored with the model version |
How to choose
Start from how the model was trained, not from which metric sounds best.
- The model's documentation says cosine, or says its vectors are unit length (most text embedding APIs): normalize on write and rank by dot product. Same order as cosine, fewer operations.
- The model was trained with a raw dot product and encodes something in the length (recommenders, some retrieval models): keep raw vectors and use the dot product. Normalizing "to be safe" throws the signal away.
- You need an ordered list (search, RAG): rank, and never threshold. A shared offset in every score leaves the order alone.
- You need a yes/no (duplicates, a cache hit, "no good answer found"): calibrate the threshold on labelled pairs from your own data, for this model, and store it with the model version.
- Whatever you pick, use one metric, one model and one normalization policy for every vector in the index. A mixture ranks wrong without an error.
What it costs
A comparison is arithmetic over the vector's dimensions:
real models produce 384 to 3,072 numbers per text, so a dot product is a few
thousand multiply-adds, and cosine adds two square roots and a division
unless the lengths are precomputed. Exact search is one comparison per
stored vector; the sentence-transformers documentation puts the practical
limit of that brute-force loop at about a million entries, after which you
move to an approximate index (primer.ml.embeddings.ann). Storage is 4 bytes
per dimension as float32 (pgvector charges 4 × dimensions + 8 bytes per
vector, half that as 16-bit floats), which primer.ml.embeddings.compression
shrinks. Normalizing is one pass at write time, once. Calibration costs
labelling: a few hundred pairs, scored and swept, and the demo shows the
return on it: a guessed threshold of 0.80 reaches F1 0.868 on the demo's
pairs, the calibrated 0.776 reaches 0.992.
What breaks
- Mixed normalization. Some vectors normalized, some not, and the ranking quietly changes: in the demo, raw dot product and cosine order 200 documents differently, and agree only after every vector is normalized. Normalize in one place, on the write path.
- A threshold from folklore. A cut-off copied from a blog post or from the previous model is meaningless: F1 swings from 0.99 to 0.87 across four hundredths of cosine. Measure it on labelled pairs, per model.
- Every score looks like 0.8. Transformer embeddings crowd into a narrow cone (Ethayarajh, 2019, found them "not isotropic in any layer"). Rank instead of judging absolute scores, or mean-centre: after centring, the demo's unrelated pairs sit at 0.00 and paraphrases at 0.30.
- A model upgrade with an old index. A new model is a new space: old
vectors, old thresholds and old scores don't carry over. Re-embed the whole
collection and recalibrate (
primer.ml.embeddings.operations). - The operator is a distance, not a similarity. Databases often expose cosine distance (1 − cosine) and the negative inner product so that ascending order is best-first, as pgvector does. Read the sign before you sort.
- Queries and documents embedded differently. Some providers train
separate treatments for the query and for the passage (Cohere's
input_typeof search_query and search_document); embed both sides the way the model expects or the scores are off. - Nearest means little on unstructured data. With 1,000 random points, the nearest is 0.7% as far as the farthest in 2 dimensions and 90% as far in 1,000: the curse of dimensionality. Real embeddings are structured, so search works, but a score gap that small is a warning that the vectors carry little.
In the wild
pgvector exposes one operator per metric on a Postgres
column: <-> for Euclidean distance, <#> for the negative inner product,
<=> for cosine distance, plus Hamming and Jaccard on bit vectors, up to
16,000 dimensions. sentence-transformers' semantic search defaults to cosine
and offers dot-product scoring for normalized vectors. OpenAI's embedding
guide states that its embeddings are normalized to length 1, recommends
cosine, and notes that it can then be computed as a plain dot product; Cohere
asks for an input_type per side of the search. Faiss's flat indexes
(IndexFlatIP, IndexFlatL2) are the exact inner-product and Euclidean
searches every approximate index is measured against. Ethayarajh (2019)
measured anisotropy in contextual models and Su et al. (2021) showed that
centring and whitening spreads the vectors back out.
Go deeper
Level 2 computes all three metrics by hand on three 3-dimensional vectors, proves in one line why normalizing makes them agree (the squared distance is 2 − 2 × cosine), builds the popularity example where cosine and dot product disagree on purpose, measures the curse of dimensionality with random points, and manufactures anisotropic embeddings to show a calibrated threshold beating a guessed one. If you only needed to choose, you are done.
Level 2
How it works, from scratch
An embedding gives every piece of text a set of coordinates on a map of meaning. Texts about similar things live on the same street; unrelated ones are across town. Draw an arrow from the centre of the map to each text's spot, and "how similar are these two texts?" becomes "how alike are these two arrows?".
There are three natural ways to compare two arrows:
- Cosine similarity compares the direction they point and ignores how long they are, like two people pointing at the same mountain, one with a longer arm.
- Euclidean distance is the straight-line distance between the two arrow tips, like measuring with a ruler.
- Dot product mixes both: pointing the same way and being long both make it bigger.
Chapter 1
A tiny worked example
Three texts, each placed in a 3-dimensional space (real models use 384 to 3,072 dimensions, meaning numbers per vector; the arithmetic is the same):
a = (1, 2, 2), b = (2, 1, 2), c = (2, −1, −2)
- Dot product: multiply position by position, then add. a·b = 1·2 + 2·1 + 2·2 = 8. a·c = 1·2 + 2·(−1) + 2·(−2) = −4.
- Length (the norm, written ‖a‖): square each number, add, and take the square root. ‖a‖ = √(1 + 4 + 4) = 3, and b and c also have length 3.
- Cosine: dot product divided by both lengths. cos(a, b) = 8 / (3·3) = 0.89. cos(a, c) = −4 / 9 = −0.44.
- Euclidean distance: subtract, square, add, square-root. ‖a − b‖ = √(1 + 1 + 0) = 1.41. ‖a − c‖ = √(1 + 9 + 16) = 5.10.
| Pair | Dot product | Lengths multiplied | Cosine | Distance | Reading |
|---|---|---|---|---|---|
| a, b | 8 | 9 | 0.89 | 1.41 | Very similar |
| a, c | −4 | 9 | −0.44 | 5.10 | Pointing away |
Cosine runs from −1 (opposite directions) through 0 (unrelated, at right angles) to 1 (same direction).
Figure 1 · Diagram
flowchart LR A["a = (1, 2, 2)"] --> D["dot product<br/>a·b = 8"] B["b = (2, 1, 2)"] --> D A --> NA["length ‖a‖ = 3"] B --> NB["length ‖b‖ = 3"] D --> COS["cosine = 8 / (3 × 3) = 0.89"] NA --> COS NB --> COS A --> SUB["difference a − b = (−1, 1, 0)"] B --> SUB SUB --> L2["distance = √2 = 1.41"]
In code: worked_example computes every number in the table above from
a, b and c, so you can check your hand arithmetic against it.
Chapter 2
The math
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| a, b | the two embedding vectors being compared | d numbers each |
| d | number of dimensions | 3 in the example; 384 to 3,072 in real models |
| i | which position (dimension) we're looking at | 1 to d |
| aᵢ | the i-th number in a | a real number |
| Σ | "add up", over i from 1 to d | |
| a · b | dot product: multiply matching positions and add | one number, any sign |
| ‖a‖ | length (norm) of a: √(Σ aᵢ²) | one number ≥ 0 |
| √ | square root | |
| cos(a, b) | cosine of the angle between a and b | −1 to 1 |
| ‖a − b‖ | Euclidean distance between the tips | ≥ 0; 0 means identical |
In words: the cosine of a and b is their dot product divided by the product of their lengths; the dot product adds up the products of matching positions; the distance is the square root of the summed squared differences.
On the example: cos(a, b) = (1·2 + 2·1 + 2·2) / (3 · 3) = 8 / 9 = 0.89; ‖a − b‖ = √((1−2)² + (2−1)² + (2−2)²) = √2 = 1.41.
In Python:
import math
a, b = [1, 2, 2], [2, 1, 2]
# a · b = Σ a_i b_i
dot = sum(a_i * b_i for a_i, b_i in zip(a, b))
dot # → 8
# ‖a‖
norm_a = math.sqrt(sum(a_i ** 2 for a_i in a))
# ‖b‖
norm_b = math.sqrt(sum(b_i ** 2 for b_i in b))
norm_a, norm_b # → (3.0, 3.0)
# cos(a, b)
round(dot / (norm_a * norm_b), 2) # → 0.89
# ‖a - b‖
round(math.sqrt(sum((a_i - b_i) ** 2 for a_i, b_i in zip(a, b))), 2) # → 1.41
The code is cosine, dot and euclidean, each a line of NumPy.
In code: rank sorts a whole collection of documents from best to worst
match under whichever of the three metrics you name.
Why it matters every vector database, semantic search box and RAG system does exactly this comparison, millions of times per second. Knowing which of the three a system uses, and why, is the difference between correct and silently wrong rankings.
Chapter 3
Normalize, and the three agree
Everyday picture trim every arrow to the same length, one unit. Now the only thing that can differ is direction, so "tips are close together" and "arrows point the same way" become the same statement.
Tiny example a = (1, 0) and b = (0.6, 0.8) both have length 1 (check: 0.6² + 0.8² = 0.36 + 0.64 = 1). Their cosine is 1·0.6 + 0·0.8 = 0.6. The squared distance is (1 − 0.6)² + (0 − 0.8)² = 0.16 + 0.64 = 0.8, and 2 − 2·0.6 = 0.8. Same number.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| ‖a − b‖² | squared distance between the tips | 0 to 4 for unit vectors |
| ‖a‖², ‖b‖² | squared lengths, both 1 after normalizing | 1 |
| a · b | dot product, equal to cosine for unit vectors | −1 to 1 |
| cos(a, b) | cosine similarity | −1 to 1 |
In words: for unit-length vectors, the squared distance is two minus twice the cosine, so the higher the cosine, the smaller the distance, always.
On the example: ‖a − b‖² = 1 + 1 − 2·0.6 = 0.8 = 2 − 2·0.6.
In Python:
# both length 1
a, b = [1, 0], [0.6, 0.8]
# for unit vectors, a · b is the cosine
cos = sum(a_i * b_i for a_i, b_i in zip(a, b))
# ‖a - b‖²
dist_sq = sum((a_i - b_i) ** 2 for a_i, b_i in zip(a, b))
# the two sides of the identity
round(dist_sq, 2), round(2 - 2 * cos, 2) # → (0.8, 0.8)
L2-normalizing (dividing a vector by its own length) is what trims every arrow to length 1. After it, the dot product is the cosine, and distance is a decreasing function of the cosine. So once vectors are L2-normalized, cosine, dot product and Euclidean distance give the same ranking.
Figure 4 · Interactive · computed from the lesson's code
Cosine similarity of two arrows
Try it: pick "Same direction, different lengths", then drag B's length. The dot product and the distance change, but the cosine stays at 1, because it only sees the angle. Drag an angle instead and watch the cosine slide from 1 through 0 to −1. Then switch on "Normalize to length 1": both arrows shrink to the unit circle, the dot product becomes the cosine, and the squared distance becomes 2 − 2 × cosine.
Figure 2 · Drawn from the lesson's code
All 400 random unit-vector pairs fall exactly on the line squared distance = 2 - 2 x cosine, so higher cosine always means smaller distance
Figure 3 · Diagram
flowchart TD
S{What was the model<br/>trained with?} -->|cosine| N[L2-normalize every vector<br/>once, at write time]
S -->|dot product| D{Does the vector length<br/>carry signal?}
D -->|yes, e.g. popularity| RAW[Keep raw vectors,<br/>rank by dot product]
D -->|no| N
N --> FAST[Rank by dot product:<br/>same order as cosine and L2,<br/>and the cheapest to compute]
In code: l2_normalize trims every row to length 1;
rankings_agree_after_normalizing ranks random documents by all three metrics
before and after normalizing and reports which rankings match, and
l2_identity_check returns both sides of the identity for two random vectors.
Why it matters vector databases typically normalize on insert and then use the dot product, the cheapest of the three. If you mix normalized and unnormalized vectors, rankings quietly change. Use the metric the embedding model was trained with.
Chapter 4
When length carries signal
Everyday picture on the map of meaning, some shops are big and popular. A recommender may deliberately make popular items' arrows longer, so they win ties. Cosine throws that away; the dot product keeps it.
Tiny example (popularity_in_the_norm): the query points along
(1, 0). A niche item (0.99, 0.10) is almost perfectly aligned. A popular item
(0.95, 0.30) × 3 = (2.85, 0.90) is slightly off but three times longer.
| Item | Cosine with query | Dot with query |
|---|---|---|
| niche (0.99, 0.10) | 0.99 / 0.995 = 0.995 | 0.99 |
| popular (2.85, 0.90) | 2.85 / 2.99 = 0.954 | 2.85 |
Cosine prefers the niche item; the dot product prefers the popular one.
Figure 5 · Drawn from the lesson's code
The short item points almost along the query and wins on cosine (0.99 vs 0.95); the item three times longer wins on dot product (2.85 vs 0.99)
Why it matters some retrieval and recommender models are trained with the raw dot product and put useful signal in the length. Normalizing their vectors "to be safe" throws that signal away.
Chapter 5
The curse of dimensionality
Everyday picture in a small village some neighbours live next door and some live across the valley. Now imagine a city with thousands of streets running in thousands of directions: pick random strangers and they all turn out to live about equally far from you. "Nearest" stops meaning much.
Tiny example scatter 1,000 random points in a cube and measure from a random spot. In 2 dimensions the nearest point is about 1/100th as far away as the farthest. In 1,000 dimensions it's about 90% as far: everyone is nearly equidistant.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| q | the query point | d numbers |
| xⱼ | the j-th of the random points | d numbers each |
| j | which random point | 1 to 1,000 |
| min, max over j | the smallest / largest value across all the points | |
| ratio | nearest distance divided by farthest distance | 0 to 1; near 1 means "all equally far" |
In words: the ratio is the distance to the closest point divided by the distance to the farthest point.
On the example: by hand first: a query at (0, 0) and three points at (1, 0), (0, 2) and (3, 4) are 1, 2 and 5 away, so the ratio is 1 / 5 = 0.2. With 1,000 random points the demo table shows about 0.01 at d = 2 and about 0.9 at d = 1,000. The Python below draws its own random points, so it lands close by rather than on the same digits: 0.02 and 0.9.
In Python:
import math, random
def ratio(q, xs):
# ‖q - x_j‖ for every j
dists = [math.dist(q, x_j) for x_j in xs]
# min over j ÷ max over j
return min(dists) / max(dists)
ratio((0, 0), [(1, 0), (0, 2), (3, 4)]) # → 0.2
random.seed(0)
for d in (2, 1000):
q = [random.random() for _ in range(d)]
xs = [[random.random() for _ in range(d)] for _ in range(1000)]
# one significant figure
print(d, f"{ratio(q, xs):.1g}") # → 2 0.02 1000 0.9
Figure 6 · Drawn from the lesson's code
The nearest-to-farthest distance ratio rises from about 0 in 2 dimensions to 0.8 at 200 and 0.92 at 2,000: random points become almost equidistant
In code: curse_of_dimensionality scatters the random points, measures
from a random query, and returns the nearest, farthest and ratio for each
dimension.
Why it matters exact nearest-neighbour search in high dimensions gets
slow and fragile, which is one reason vector search uses approximate indexes
(primer.ml.embeddings.ann). Real embeddings aren't random: meaning puts
them on a much lower-dimensional structure, so nearest-neighbour search still
works well in practice.
Chapter 6
Anisotropy: why every score looks like 0.8
Everyday picture a stadium crowd all facing the stage. Ask "are these two people facing the same way?" and the answer is "roughly yes" for everyone, so the question tells you little. Many embedding models behave like that crowd: their vectors bunch in a narrow cone, so even unrelated texts get cosines of 0.7 or more. This lopsidedness is called anisotropy ("not the same in every direction").
Tiny example build each vector as √3·m + n, where m is one direction everybody shares (the stage) and n is a unit-length direction of each text's own. For two unrelated texts the n parts are at right angles, so their dot product is (√3)² = 3 and each length is √(3 + 1) = 2, and the cosine is 3 / (2 · 2) = 0.75. For paraphrases, whose own parts agree with cosine ρ, the cosine becomes 0.75 + 0.25·ρ: about 0.79 to 0.86. All the signal is squeezed into a few hundredths.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape / range |
|---|---|---|
| v, v₁, v₂ | an embedding vector | d numbers (768 here) |
| m | the shared direction every vector leans toward | unit vector |
| n | the text's own direction | unit vector |
| α (alpha) | how strongly every vector leans toward m | √3 here |
| β (beta) | how strongly each keeps its own direction | 1 here |
| ρ (rho) | cosine between the two texts' own directions | 0 unrelated; 0.15 to 0.45 for paraphrases |
In words: every vector is a big shared part plus a small personal part, so the cosine is mostly the shared part's weight, nudged up a little by how much the personal parts agree.
On the example: unrelated, ρ = 0: (3 + 0) / (3 + 1) = 0.75. A paraphrase with ρ = 0.4: (3 + 0.4) / 4 = 0.85.
In Python:
import math
alpha, beta = math.sqrt(3), 1
# shared direction; two unrelated own directions
m, n1, n2 = (1, 0, 0), (0, 1, 0), (0, 0, 1)
# v = α m + β n
v1 = [alpha * m_i + beta * n_i for m_i, n_i in zip(m, n1)]
v2 = [alpha * m_i + beta * n_i for m_i, n_i in zip(m, n2)]
dot = sum(a * b for a, b in zip(v1, v2))
# the cosine, measured directly
round(dot / (math.hypot(*v1) * math.hypot(*v2)), 2) # → 0.75
def cos_formula(rho):
return (alpha ** 2 + beta ** 2 * rho) / (alpha ** 2 + beta ** 2)
# unrelated, and a paraphrase
round(cos_formula(0), 2), round(cos_formula(0.4), 2) # → (0.75, 0.85)
Figure 7 · Drawn from the lesson's code
Raw cosines crowd paraphrases (0.76 to 0.87) and unrelated pairs (0.72 to 0.78) together; mean-centred, unrelated sit near 0 and paraphrases near 0.3
In code: anisotropic_embeddings builds the pairs as α·m + β·n,
pair_cosines scores each pair, and mean_center subtracts the average
vector to remove the shared direction.
Picking a threshold from data, not folklore
Sometimes you need a yes/no cut-off: "are these duplicates?", "is this cached answer close enough?". Pick it by measuring. Label some pairs, score them, and choose the threshold with the best F1, the balance of precision (of the pairs I called similar, how many were?) and recall (of the truly similar pairs, how many did I catch?):
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Range |
|---|---|---|
| P | precision: true matches ÷ everything called a match | 0 to 1 |
| R | recall: true matches found ÷ all true matches | 0 to 1 |
| F₁ | their harmonic mean; high only when both are high | 0 to 1 |
In words: F1 is twice the product of precision and recall, divided by their sum.
On a hand example: scores 0.9, 0.8, 0.7, 0.6 where the first two are true
matches. Threshold 0.8 calls exactly those two matches: P = 2/2 = 1,
R = 2/2 = 1, F₁ = 2·1·1 / 2 = 1. best_threshold finds 0.8.
In Python:
scores = [0.9, 0.8, 0.7, 0.6]
# the labels
is_match = [True, True, False, False]
# threshold 0.8
called = [s >= 0.8 for s in scores]
true_matches = sum(c and y for c, y in zip(called, is_match))
# precision
P = true_matches / sum(called)
# recall
R = true_matches / sum(is_match)
# F₁
2 * P * R / (P + R) # → 1.0
Figure 9 · Diagram
flowchart LR L[Sample a few hundred pairs<br/>from your own data] --> H[Label each:<br/>same meaning or not] H --> E[Score every pair<br/>with THIS model] E --> W[Sweep thresholds,<br/>compute precision, recall, F1] W --> P[Pick the threshold that fits<br/>your cost of each error] P --> V[Store it with the<br/>model version] V -->|model upgraded| E
Figure 8 · Drawn from the lesson's code
F1 on raw scores swings within a few hundredths: the calibrated threshold 0.776 reaches F1 0.99, while the guessed 0.80 drops to 0.87
In code: f1_curve computes F1 at every candidate threshold for this
plot; best_threshold returns the single best one with its precision and
recall.
Why it matters
- Absolute scores mean different things for different models. "0.8" is not a universal "similar".
- Rank, don't threshold, whenever you can. A shared offset doesn't change which document is on top.
- When you must threshold (dedup, semantic caching, "no good answer found"), calibrate on labeled pairs for that model, and redo it when the model changes.
Test yourself
4 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Q: When do cosine, dot product and Euclidean distance give the same ranking?Think it through, then reveal
When all vectors are L2-normalized: the dot product equals the cosine, and the squared distance equals 2 − 2cos, which always moves opposite to cosine.
Question 2Q: When would you deliberately use the raw dot product on unnormalized vectors?Think it through, then reveal
When the model was trained that way and encodes useful signal in the length (e.g. popularity or quality in some recommender and retrieval models). Normalizing would throw that signal away.
Question 3Q: Similarity scores all sit between 0.75 and 0.85. What's going on and how do you set a threshold?Think it through, then reveal
The model is anisotropic: its vectors share a dominant direction, so every pair looks similar and the real signal lives in a narrow band. Ranking still works. For a threshold, label a few hundred pairs from your own data, sweep the threshold, and pick the one that maximizes F1 (or matches your precision/recall needs). Optionally mean-center the vectors. Redo it for every model version.
Question 4Q: What is the curse of dimensionality for nearest-neighbour search?Think it through, then reveal
For random high-dimensional data, the nearest and farthest points are almost the same distance away, so "nearest" carries little information and exact search gets expensive. Real embeddings have structure, and approximate indexes trade a little recall for large speedups.
Primary sources
The papers behind this lesson
Proved that for broad families of random data, the nearest and farthest neighbours become equally far as dimension grows: the curse of dimensionality measured in this lesson.
The paper ↗Showed that transformer embeddings occupy a narrow cone (anisotropy), which is why raw cosine scores bunch together.
The paper ↗Showed that centering and whitening sentence embeddings spreads them back out and improves similarity quality.
The paper ↗Researcher's shelf
Further reading
- Cosine similarity (Wikipedia): https://en.wikipedia.org/wiki/Cosine_similarity
- Beyer et al., When Is "Nearest Neighbor" Meaningful? (1999): https://doi.org/10.1007/3-540-49257-7_15
- Ethayarajh, How Contextual are Contextualized Word Representations? (anisotropy, 2019): https://arxiv.org/abs/1909.00512
- Su et al., Whitening Sentence Representations (2021): https://arxiv.org/abs/2103.15316
- sentence-transformers, semantic search and score functions: https://www.sbert.net/examples/applications/semantic-search/README.html
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.