rumblr Work in progressWIP

● The AI Primer · Lesson 27 · Embeddings, the centerpiece

Similarity

how close are two meanings?

This lesson covers Cosine vs. dot vs. distance, normalization, anisotropy, thresholds

Members · open during launch 27 min8 figures and diagrams1 interactive
How it works builds the idea from scratch. Math & code adds the formulas and the Python.

At a glance

Key takeaways

  1. Cosine compares direction, the dot product compares direction and length, Euclidean distance measures the gap between the tips.
  2. On L2-normalized vectors all three give the same ranking, because ‖a − b‖² = 2 − 2cos.
  3. Use the metric the model was trained with.
  4. Scores aren't comparable across models (anisotropy). Rank when you can; calibrate thresholds on labeled pairs when you can't.

Level 2

How it works, from scratch

An embedding gives every piece of text a set of coordinates on a map of meaning. Texts about similar things live on the same street; unrelated ones are across town. Draw an arrow from the centre of the map to each text's spot, and "how similar are these two texts?" becomes "how alike are these two arrows?".

There are three natural ways to compare two arrows:

  • Cosine similarity compares the direction they point and ignores how long they are, like two people pointing at the same mountain, one with a longer arm.
  • Euclidean distance is the straight-line distance between the two arrow tips, like measuring with a ruler.
  • Dot product mixes both: pointing the same way and being long both make it bigger.

Chapter 1

A tiny worked example

Three texts, each placed in a 3-dimensional space (real models use 384 to 3,072 dimensions, meaning numbers per vector; the arithmetic is the same):

a = (1, 2, 2), b = (2, 1, 2), c = (2, −1, −2)

  1. Dot product: multiply position by position, then add. a·b = 1·2 + 2·1 + 2·2 = 8. a·c = 1·2 + 2·(−1) + 2·(−2) = −4.
  2. Length (the norm, written ‖a‖): square each number, add, and take the square root. ‖a‖ = √(1 + 4 + 4) = 3, and b and c also have length 3.
  3. Cosine: dot product divided by both lengths. cos(a, b) = 8 / (3·3) = 0.89. cos(a, c) = −4 / 9 = −0.44.
  4. Euclidean distance: subtract, square, add, square-root. ‖a − b‖ = √(1 + 1 + 0) = 1.41. ‖a − c‖ = √(1 + 9 + 16) = 5.10.
Pair Dot product Lengths multiplied Cosine Distance Reading
a, b 8 9 0.89 1.41 Very similar
a, c −4 9 −0.44 5.10 Pointing away

Cosine runs from −1 (opposite directions) through 0 (unrelated, at right angles) to 1 (same direction).

Figure 1 · Diagram

Reading it: the two input vectors feed three calculations. The top path is the dot product. The middle paths measure each vector's length, and cosine is simply the dot product divided by those lengths, which is how it comes to "ignore length". The bottom path subtracts the vectors and measures what's left: that's distance. Everything later in this lesson is these three boxes.

In code: worked_example computes every number in the table above from a, b and c, so you can check your hand arithmetic against it.

Chapter 2

The math

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
a, b the two embedding vectors being compared d numbers each
d number of dimensions 3 in the example; 384 to 3,072 in real models
i which position (dimension) we're looking at 1 to d
aᵢ the i-th number in a a real number
Σ "add up", over i from 1 to d
a · b dot product: multiply matching positions and add one number, any sign
‖a‖ length (norm) of a: √(Σ aᵢ²) one number ≥ 0
√ square root
cos(a, b) cosine of the angle between a and b −1 to 1
‖a − b‖ Euclidean distance between the tips ≥ 0; 0 means identical

In words: the cosine of a and b is their dot product divided by the product of their lengths; the dot product adds up the products of matching positions; the distance is the square root of the summed squared differences.

On the example: cos(a, b) = (1·2 + 2·1 + 2·2) / (3 · 3) = 8 / 9 = 0.89; ‖a − b‖ = √((1−2)² + (2−1)² + (2−2)²) = √2 = 1.41.

In Python:

import math
a, b = [1, 2, 2], [2, 1, 2]
# a · b = Σ a_i b_i
dot = sum(a_i * b_i for a_i, b_i in zip(a, b))
dot  # → 8
# ‖a‖
norm_a = math.sqrt(sum(a_i ** 2 for a_i in a))
# ‖b‖
norm_b = math.sqrt(sum(b_i ** 2 for b_i in b))
norm_a, norm_b  # → (3.0, 3.0)
# cos(a, b)
round(dot / (norm_a * norm_b), 2)  # → 0.89
# ‖a - b‖
round(math.sqrt(sum((a_i - b_i) ** 2 for a_i, b_i in zip(a, b))), 2)  # → 1.41

The code is cosine, dot and euclidean, each a line of NumPy.

In code: rank sorts a whole collection of documents from best to worst match under whichever of the three metrics you name.

Why it matters every vector database, semantic search box and RAG system does exactly this comparison, millions of times per second. Knowing which of the three a system uses, and why, is the difference between correct and silently wrong rankings.

Chapter 3

Normalize, and the three agree

Everyday picture trim every arrow to the same length, one unit. Now the only thing that can differ is direction, so "tips are close together" and "arrows point the same way" become the same statement.

Tiny example a = (1, 0) and b = (0.6, 0.8) both have length 1 (check: 0.6² + 0.8² = 0.36 + 0.64 = 1). Their cosine is 1·0.6 + 0·0.8 = 0.6. The squared distance is (1 − 0.6)² + (0 − 0.8)² = 0.16 + 0.64 = 0.8, and 2 − 2·0.6 = 0.8. Same number.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
‖a − b‖² squared distance between the tips 0 to 4 for unit vectors
‖a‖², ‖b‖² squared lengths, both 1 after normalizing 1
a · b dot product, equal to cosine for unit vectors −1 to 1
cos(a, b) cosine similarity −1 to 1

In words: for unit-length vectors, the squared distance is two minus twice the cosine, so the higher the cosine, the smaller the distance, always.

On the example: ‖a − b‖² = 1 + 1 − 2·0.6 = 0.8 = 2 − 2·0.6.

In Python:

# both length 1
a, b = [1, 0], [0.6, 0.8]
# for unit vectors, a · b is the cosine
cos = sum(a_i * b_i for a_i, b_i in zip(a, b))
# ‖a - b‖²
dist_sq = sum((a_i - b_i) ** 2 for a_i, b_i in zip(a, b))
# the two sides of the identity
round(dist_sq, 2), round(2 - 2 * cos, 2)  # → (0.8, 0.8)

L2-normalizing (dividing a vector by its own length) is what trims every arrow to length 1. After it, the dot product is the cosine, and distance is a decreasing function of the cosine. So once vectors are L2-normalized, cosine, dot product and Euclidean distance give the same ranking.

Figure 2 · Interactive · computed from the lesson's code

Cosine similarity of two arrows

Try it: pick "Same direction, different lengths", then drag B's length. The dot product and the distance change, but the cosine stays at 1, because it only sees the angle. Drag an angle instead and watch the cosine slide from 1 through 0 to −1. Then switch on "Normalize to length 1": both arrows shrink to the unit circle, the dot product becomes the cosine, and the squared distance becomes 2 − 2 × cosine.

Figure 3 · Drawn from the lesson's code

−1.00 −0.75 −0.50 −0.25 0.00 0.25 0.50 0.75 1.00 cosine similarity 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 4.0 squared Euclidean distance On unit vectors, distance is a function of cosine y = 2 - 2x random unit-vector pairs

All 400 random unit-vector pairs fall exactly on the line squared distance = 2 - 2 x cosine, so higher cosine always means smaller distance

Reading it: each dot is one pair of random unit vectors. The horizontal axis is their cosine and the vertical axis is their squared Euclidean distance. Every dot sits exactly on the line y = 2 − 2x. Because the line only goes down, "higher cosine" and "smaller distance" always mean the same thing, so the two metrics can never disagree about which neighbour is nearer.

Figure 4 · Diagram

Reading it: start at the top with the question that matters, which is how the model was trained, not which metric sounds best. Most text embedding models are trained with cosine, so you normalize once when you write vectors, then use the plain dot product at query time. It gives the same order as cosine and costs less. Only when length is deliberate signal do you skip normalizing.

In code: l2_normalize trims every row to length 1; rankings_agree_after_normalizing ranks random documents by all three metrics before and after normalizing and reports which rankings match, and l2_identity_check returns both sides of the identity for two random vectors.

Why it matters vector databases typically normalize on insert and then use the dot product, the cheapest of the three. If you mix normalized and unnormalized vectors, rankings quietly change. Use the metric the embedding model was trained with.

Chapter 4

When length carries signal

Everyday picture on the map of meaning, some shops are big and popular. A recommender may deliberately make popular items' arrows longer, so they win ties. Cosine throws that away; the dot product keeps it.

Tiny example (popularity_in_the_norm): the query points along (1, 0). A niche item (0.99, 0.10) is almost perfectly aligned. A popular item (0.95, 0.30) × 3 = (2.85, 0.90) is slightly off but three times longer.

Item Cosine with query Dot with query
niche (0.99, 0.10) 0.99 / 0.995 = 0.995 0.99
popular (2.85, 0.90) 2.85 / 2.99 = 0.954 2.85

Cosine prefers the niche item; the dot product prefers the popular one.

Figure 5 · Drawn from the lesson's code

0.0 0.5 1.0 1.5 2.0 2.5 3.0 dimension 1 0.00 0.25 0.50 0.75 1.00 dimension 2 Cosine picks blue (angle); dot product picks orange (length) query niche item (aligned, short) popular item (off-angle, long)

The short item points almost along the query and wins on cosine (0.99 vs 0.95); the item three times longer wins on dot product (2.85 vs 0.99)

Reading it: the black arrow is the query. The blue arrow (niche item) points almost exactly the same way but is short. The orange arrow (popular item) leans a little away but is three times longer. Cosine only looks at the angle between each arrow and the query, so blue wins. The dot product also rewards length, so orange wins. Neither is wrong: it depends on what the model was trained to encode.

Why it matters some retrieval and recommender models are trained with the raw dot product and put useful signal in the length. Normalizing their vectors "to be safe" throws that signal away.

Chapter 5

The curse of dimensionality

Everyday picture in a small village some neighbours live next door and some live across the valley. Now imagine a city with thousands of streets running in thousands of directions: pick random strangers and they all turn out to live about equally far from you. "Nearest" stops meaning much.

Tiny example scatter 1,000 random points in a cube and measure from a random spot. In 2 dimensions the nearest point is about 1/100th as far away as the farthest. In 1,000 dimensions it's about 90% as far: everyone is nearly equidistant.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
q the query point d numbers
xⱼ the j-th of the random points d numbers each
j which random point 1 to 1,000
min, max over j the smallest / largest value across all the points
ratio nearest distance divided by farthest distance 0 to 1; near 1 means "all equally far"

In words: the ratio is the distance to the closest point divided by the distance to the farthest point.

On the example: by hand first: a query at (0, 0) and three points at (1, 0), (0, 2) and (3, 4) are 1, 2 and 5 away, so the ratio is 1 / 5 = 0.2. With 1,000 random points the demo table shows about 0.01 at d = 2 and about 0.9 at d = 1,000. The Python below draws its own random points, so it lands close by rather than on the same digits: 0.02 and 0.9.

In Python:

import math, random
def ratio(q, xs):
    # ‖q - x_j‖ for every j
    dists = [math.dist(q, x_j) for x_j in xs]
    # min over j ÷ max over j
    return min(dists) / max(dists)
ratio((0, 0), [(1, 0), (0, 2), (3, 4)])  # → 0.2
random.seed(0)
for d in (2, 1000):
    q = [random.random() for _ in range(d)]
    xs = [[random.random() for _ in range(d)] for _ in range(1000)]
    # one significant figure
    print(d, f"{ratio(q, xs):.1g}")  # → 2 0.02 1000 0.9

Figure 6 · Drawn from the lesson's code

1 0 0 1 0 1 1 0 2 1 0 3 dimensions (log scale) 0.0 0.2 0.4 0.6 0.8 1.0 nearest distance / farthest distance Random points become equidistant as dimension grows

The nearest-to-farthest distance ratio rises from about 0 in 2 dimensions to 0.8 at 200 and 0.92 at 2,000: random points become almost equidistant

Reading it: the horizontal axis is the number of dimensions (log scale); the vertical axis is the nearest-to-farthest ratio above. In 2-D the ratio is near 0: the nearest point really is much nearer. By a few hundred dimensions the curve flattens toward 1, so every random point is roughly equally far away, and "nearest" tells you almost nothing about random data.

In code: curse_of_dimensionality scatters the random points, measures from a random query, and returns the nearest, farthest and ratio for each dimension.

Why it matters exact nearest-neighbour search in high dimensions gets slow and fragile, which is one reason vector search uses approximate indexes (primer.ml.embeddings.ann). Real embeddings aren't random: meaning puts them on a much lower-dimensional structure, so nearest-neighbour search still works well in practice.

Chapter 6

Anisotropy: why every score looks like 0.8

Everyday picture a stadium crowd all facing the stage. Ask "are these two people facing the same way?" and the answer is "roughly yes" for everyone, so the question tells you little. Many embedding models behave like that crowd: their vectors bunch in a narrow cone, so even unrelated texts get cosines of 0.7 or more. This lopsidedness is called anisotropy ("not the same in every direction").

Tiny example build each vector as √3·m + n, where m is one direction everybody shares (the stage) and n is a unit-length direction of each text's own. For two unrelated texts the n parts are at right angles, so their dot product is (√3)² = 3 and each length is √(3 + 1) = 2, and the cosine is 3 / (2 · 2) = 0.75. For paraphrases, whose own parts agree with cosine ρ, the cosine becomes 0.75 + 0.25·ρ: about 0.79 to 0.86. All the signal is squeezed into a few hundredths.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
v, v₁, v₂ an embedding vector d numbers (768 here)
m the shared direction every vector leans toward unit vector
n the text's own direction unit vector
α (alpha) how strongly every vector leans toward m √3 here
β (beta) how strongly each keeps its own direction 1 here
ρ (rho) cosine between the two texts' own directions 0 unrelated; 0.15 to 0.45 for paraphrases

In words: every vector is a big shared part plus a small personal part, so the cosine is mostly the shared part's weight, nudged up a little by how much the personal parts agree.

On the example: unrelated, ρ = 0: (3 + 0) / (3 + 1) = 0.75. A paraphrase with ρ = 0.4: (3 + 0.4) / 4 = 0.85.

In Python:

import math
alpha, beta = math.sqrt(3), 1
# shared direction; two unrelated own directions
m, n1, n2 = (1, 0, 0), (0, 1, 0), (0, 0, 1)
# v = α m + β n
v1 = [alpha * m_i + beta * n_i for m_i, n_i in zip(m, n1)]
v2 = [alpha * m_i + beta * n_i for m_i, n_i in zip(m, n2)]
dot = sum(a * b for a, b in zip(v1, v2))
# the cosine, measured directly
round(dot / (math.hypot(*v1) * math.hypot(*v2)), 2)  # → 0.75
def cos_formula(rho):
    return (alpha ** 2 + beta ** 2 * rho) / (alpha ** 2 + beta ** 2)
# unrelated, and a paraphrase
round(cos_formula(0), 2), round(cos_formula(0.4), 2)  # → (0.75, 0.85)

Figure 7 · Drawn from the lesson's code

0.72 0.74 0.76 0.78 0.80 0.82 0.84 0.86 0.88 cosine similarity 0 10 20 30 40 number of pairs raw cosine: everything is 0.7 to 0.9 paraphrase pairs unrelated pairs −0.1 0.0 0.1 0.2 0.3 0.4 0.5 cosine similarity 0 10 20 30 40 50 number of pairs after mean-centering: the gap opens paraphrase pairs unrelated pairs

Raw cosines crowd paraphrases (0.76 to 0.87) and unrelated pairs (0.72 to 0.78) together; mean-centred, unrelated sit near 0 and paraphrases near 0.3

Reading it: the left panel shows raw cosine scores for 300 paraphrase pairs (one colour) and 300 unrelated pairs (the other). Both piles sit between 0.7 and 0.9, so a score of "0.8" alone says little. The right panel shows the same pairs after mean-centering, which means subtracting the average vector of the whole collection (that removes the shared "stage" direction). Unrelated pairs drop to about 0, paraphrases sit clearly above, and the gap is easy to see. The information was there all along, squeezed into a narrow band.

In code: anisotropic_embeddings builds the pairs as α·m + β·n, pair_cosines scores each pair, and mean_center subtracts the average vector to remove the shared direction.

Picking a threshold from data, not folklore

Sometimes you need a yes/no cut-off: "are these duplicates?", "is this cached answer close enough?". Pick it by measuring. Label some pairs, score them, and choose the threshold with the best F1, the balance of precision (of the pairs I called similar, how many were?) and recall (of the truly similar pairs, how many did I catch?):

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Range
P precision: true matches ÷ everything called a match 0 to 1
R recall: true matches found ÷ all true matches 0 to 1
F₁ their harmonic mean; high only when both are high 0 to 1

In words: F1 is twice the product of precision and recall, divided by their sum.

On a hand example: scores 0.9, 0.8, 0.7, 0.6 where the first two are true matches. Threshold 0.8 calls exactly those two matches: P = 2/2 = 1, R = 2/2 = 1, F₁ = 2·1·1 / 2 = 1. best_threshold finds 0.8.

In Python:

scores = [0.9, 0.8, 0.7, 0.6]
# the labels
is_match = [True, True, False, False]
# threshold 0.8
called = [s >= 0.8 for s in scores]
true_matches = sum(c and y for c, y in zip(called, is_match))
# precision
P = true_matches / sum(called)
# recall
R = true_matches / sum(is_match)
# F₁
2 * P * R / (P + R)  # → 1.0

Figure 8 · Diagram

Reading it: a threshold is a property of one model on one kind of data, so it is measured, not guessed. Label pairs, score them, sweep, and pick. The loop back from "model upgraded" is the step people forget: a new model has a different score distribution, so the old threshold is meaningless.

Figure 9 · Drawn from the lesson's code

0.72 0.74 0.76 0.78 0.80 0.82 0.84 0.86 0.88 threshold on raw cosine 0.0 0.2 0.4 0.6 0.8 1.0 F1 Threshold choice on an anisotropic model guessed 0.80 calibrated 0.776

F1 on raw scores swings within a few hundredths: the calibrated threshold 0.776 reaches F1 0.99, while the guessed 0.80 drops to 0.87

Reading it: the curve is F1 at each candidate threshold on the raw anisotropic scores. The dashed line is the "obvious" guess of 0.8; the solid line is the calibrated threshold. F1 changes sharply over a few hundredths of cosine, which is exactly why a threshold carried over from a different model (or a blog post) can be badly wrong.

In code: f1_curve computes F1 at every candidate threshold for this plot; best_threshold returns the single best one with its precision and recall.

Why it matters

  • Absolute scores mean different things for different models. "0.8" is not a universal "similar".
  • Rank, don't threshold, whenever you can. A shared offset doesn't change which document is on top.
  • When you must threshold (dedup, semantic caching, "no good answer found"), calibrate on labeled pairs for that model, and redo it when the model changes.

Test yourself

4 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: When do cosine, dot product and Euclidean distance give the same ranking?Think it through, then reveal

When all vectors are L2-normalized: the dot product equals the cosine, and the squared distance equals 2 − 2cos, which always moves opposite to cosine.

Question 2Q: When would you deliberately use the raw dot product on unnormalized vectors?Think it through, then reveal

When the model was trained that way and encodes useful signal in the length (e.g. popularity or quality in some recommender and retrieval models). Normalizing would throw that signal away.

Question 3Q: Similarity scores all sit between 0.75 and 0.85. What's going on and how do you set a threshold?Think it through, then reveal

The model is anisotropic: its vectors share a dominant direction, so every pair looks similar and the real signal lives in a narrow band. Ranking still works. For a threshold, label a few hundred pairs from your own data, sweep the threshold, and pick the one that maximizes F1 (or matches your precision/recall needs). Optionally mean-center the vectors. Redo it for every model version.

Question 4Q: What is the curse of dimensionality for nearest-neighbour search?Think it through, then reveal

For random high-dimensional data, the nearest and farthest points are almost the same distance away, so "nearest" carries little information and exact search gets expensive. Real embeddings have structure, and approximate indexes trade a little recall for large speedups.

Primary sources

The papers behind this lesson

Beyer, Goldstein, Ramakrishnan and Shaft, When Is "Nearest Neighbor" Meaningful? (1999)

Proved that for broad families of random data, the nearest and farthest neighbours become equally far as dimension grows: the curse of dimensionality measured in this lesson.

The paper ↗
Ethayarajh, How Contextual are Contextualized Word Representations? (2019)

Showed that transformer embeddings occupy a narrow cone (anisotropy), which is why raw cosine scores bunch together.

The paper ↗
Su, Cao, Liu and Ou, Whitening Sentence Representations for Better Semantics and Faster Retrieval (2021)

Showed that centering and whitening sentence embeddings spreads them back out and improves similarity quality.

The paper ↗

Researcher's shelf

Further reading

  • Cosine similarity (Wikipedia): https://en.wikipedia.org/wiki/Cosine_similarity
  • Beyer et al., When Is "Nearest Neighbor" Meaningful? (1999): https://doi.org/10.1007/3-540-49257-7_15
  • Ethayarajh, How Contextual are Contextualized Word Representations? (anisotropy, 2019): https://arxiv.org/abs/1909.00512
  • Su et al., Whitening Sentence Representations (2021): https://arxiv.org/abs/2103.15316
  • sentence-transformers, semantic search and score functions: https://www.sbert.net/examples/applications/semantic-search/README.html

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.