rumblr Work in progressWIP

● The AI Primer · Lesson 26 · Embeddings, the centerpiece

word2vec

learning meaning from the company a word keeps

You'll be able to explain Where embeddings came from, analogies, the "bank" problem

Members · open during launch 29 min6 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. word2vec learns a vector per word by guessing nearby words; words that keep the same company get similar vectors.
  2. Negative sampling makes it cheap: raise the score of the true pair, lower the score of k random pairs, instead of scoring the whole vocabulary.
  3. Consistent differences in context become consistent directions, hence king − man + woman ≈ queen.
  4. GloVe and PPMI + SVD count co-occurrences and compress them, and reach the same geometry.
  5. Static vectors can't separate word senses; contextual (transformer) embeddings can, and they replaced static ones.

Level 1

The practitioner's guide

In one sentence

word2vec turns every word in a body of text into a short list of numbers (its vector) by training it to guess its neighbours, so that words used in the same company land close together and can be compared, averaged and searched with arithmetic.

When you need it

You need word vectors the moment a program has to know that two different strings mean similar things: matching "laptop" to "notebook computer" in a search box, grouping support tickets by topic, feeding words to a classifier as something richer than an id, or finding which products are "like" a given one from the sentences they appear in. The tell is a synonym table you maintain by hand, or a keyword search that misses every rephrasing. You don't need word2vec when a word's meaning depends on the sentence around it: in this lesson's demo the single static vector for "bank" sits at cosine 0.52 to the river words and 0.50 to the money words, halfway between its two senses, where one step of attention over "river bank fish" takes it to 0.83 on the river side and "bank loan cash" to 0.85 on the money side. And you don't need it for whole sentences or documents: averaging word vectors is a rough baseline, but the sentence-embedding models of primer.ml.embeddings.contrastive were built for that job. Today word2vec is the idea to understand, and a cheap tool for a vocabulary of your own; contextual models are the default for text.

Your options

Five ways to get a vector per word, from the cheapest to the most capable:

Option What it does What it guarantees What it costs Where it lives
Download pretrained vectors (GloVe, word2vec, fastText) Loads a table trained on billions of words of news, Wikipedia or web text A good general vocabulary in minutes: Stanford's largest GloVe set covers 2.2 million words at 300 dimensions, from 840 billion tokens A file of gigabytes; words your field uses differently keep their public meaning A file you load
Count, then compress (PPMI + SVD) Tallies which words appear near which, keeps the pairs that meet more often than chance, compresses the table Deterministic, one pass over the text, no training loop to tune Memory for a vocabulary-by-vocabulary table, which caps the vocabulary Your code
Train skip-gram with negative sampling on your corpus Plays the guessing game on your own text Vectors that know your jargon; the analogies of Level 2 on clean data CPU hours in proportion to the text, and a corpus big enough to see each word in many contexts A library such as gensim, on your machine
Subword vectors (fastText) Builds each word's vector from the vectors of its character pieces A vector for words never seen in training, including typos and rare inflections A bigger model, and pieces shared by unrelated words leak into each other A library
Contextual embeddings (BERT and every model since) Runs the sentence through a transformer and reads off a fresh vector per word, per sentence Word senses separated by their context A forward pass per text, a model to host or an API to pay A model server or an embedding API

How to choose

Start from whose words they are and whether context matters.

  • General English, a prototype by this afternoon: download pretrained vectors and get on with it.
  • Your own vocabulary (product codes, ticket jargon, a legal field): train skip-gram on your own text. It runs on a laptop; the demo here trains on 1,800 sentences in seconds.
  • Typos, rare words, or a language with many word forms: fastText, which assembles a vector for any spelling.
  • A word whose meaning depends on the sentence, or whole sentences to compare: a contextual model, and the rest of this package.
  • Whatever you pick, judge the vectors on your own task, not on analogy puzzles. Level 2's corpus is built from three clean attributes so that king − man + woman lands on queen; your search logs are messier, and they are what counts.

What it costs

Training is cheap, which is the point of the method: negative sampling scores one true pair and k noise words per training example instead of the whole vocabulary. The negative-sampling paper reports k = 5 to 20 as useful for small datasets and 2 to 5 for large ones, an optimized single-machine implementation training on more than 100 billion words in a day, and a 2× to 10× further speed-up from subsampling the most frequent words. This lesson trains 16-dimensional vectors for a 45-word vocabulary in seconds. Storage is a table of vocabulary × dimensions numbers: the 400,000-word, 300-dimension GloVe set is 400,000 × 300 × 4 bytes, about 480 MB as float32. Query time costs nothing: a vector is a table lookup, no model runs. Quality, in the paper's own numbers: 300-dimensional vectors trained on a billion words score about 60% on its analogy test, and the same paper names the settings that matter most as the architecture, the vector size, the subsampling rate and the window.

What breaks

  • One vector, many senses. "bank" at 0.52 to river and 0.50 to money is the whole story: a static table cannot separate senses. Use a contextual model where senses matter.
  • Words never seen. A word absent from training has no vector, and libraries drop rare words on purpose (gensim's default keeps words seen at least 5 times). Map unknown words to a shared placeholder, or use fastText.
  • Order is invisible. The window records which words were near, not in what order: "not good" and "very good" put "good" in the same company. Sentiment and negation need a model that reads order.
  • Frequent words swamp the rest. "the" appears next to everything and teaches nothing; the 3/4 power on the noise distribution and subsampling of frequent words exist to counter it. Skip them in a home-made trainer and the vectors degrade.
  • The corpus's prejudices come along. Bolukbasi et al. (2016) found that vectors trained on Google News exhibit gender stereotypes "to a disturbing extent". Audit before any decision about people rests on them.
  • Public vectors, private meaning. A pretrained "python" is a snake and a language in whatever balance the web had; your codebase has one meaning. Train on your own text when the two differ.

In the wild

gensim's Word2Vec is the standard Python implementation (defaults: 100 dimensions, window 5, 5 negatives, 5 epochs, min_count 5, and the CBOW variant unless you ask for skip-gram). Stanford publishes GloVe vectors trained on Wikipedia and Gigaword (6 billion tokens, 400,000 words, 50 to 300 dimensions) and on Common Crawl (840 billion tokens, 2.2 million words, 300 dimensions). fastText (Bojanowski et al., 2016) is the subword variant. Levy and Goldberg (2014) showed that skip-gram with negative sampling factorizes a shifted PMI table, so counting and predicting are one family. BERT (2018) is the contextual model that replaced static tables as the default, and every sentence-embedding model since inherits its shape.

Go deeper

Level 2 plays the guessing game by hand on a three-word sentence: the pairs a window makes, the dot product, the sigmoid, one nudge of the vectors, and the gradient that nudge follows. It then trains real vectors on a small corpus, checks king − man + woman = queen, gets the same vectors by counting (PPMI + SVD, GloVe), and shows one attention step turning a static "bank" into a contextual one. If you only needed to choose, you are done.

Level 2

How it works, from scratch

You can learn a lot about a stranger from their friends. If two people keep turning up with the same crowd, they probably have something in common. Linguists put it the same way: "You shall know a word by the company it keeps" (J. R. Firth, 1957). "Coffee" and "tea" both show up next to "cup", "hot" and "morning", so they must be related.

word2vec turns that idea into a game. Every word gets a spot on a map of meaning (a short list of numbers, its vector). The game: looking only at a word's spot, guess which words sit near it in real sentences. Every wrong guess nudges the spots a little. After millions of nudges, words that keep the same company end up in the same neighbourhood, because that's the only way to guess well for all of them.

Chapter 1

A tiny worked example: one sentence, one nudge

Step 1, make guessing pairs. Take the sentence "river bank fish" and a window of 1 (look one word left and right). Each word, the center, is paired with each neighbour, its context:

(river, bank), (bank, river), (bank, fish), (fish, bank)

No human labels anything: the text itself says which words appeared together.

Step 2, score a pair. Put everything in 2 dimensions. The center word has vector v = (1, 0), its true context word has u = (0, 1), and one random "noise" word (a word we pretend is a context, to have something to push away) has uₙ = (0, −1). The score of a pair is their dot product (multiply matching positions, add): v·u = 1·0 + 0·1 = 0, and v·uₙ = 0. Zero means "no opinion".

Step 3, turn scores into probabilities with the sigmoid σ(x), which squashes any number into 0 to 1 (σ(0) = 0.5, big positive → near 1, big negative → near 0). The model says the true pair is 50% likely and the noise pair is 50% likely. That's bad: we want ~100% and ~0%.

Step 4, measure the mistake as a loss (a single number that's big when the guesses are bad): −ln σ(0) − ln σ(−0) = 0.693 + 0.693 = 1.386.

Step 5, nudge. Move v a little toward u and away from uₙ, and move u and uₙ a little toward or away from v. With a step size (learning rate) of 0.1 we get v = (1, 0.1), u = (0.05, 1), uₙ = (−0.05, −1). Now v·u = 0.15 and v·uₙ = −0.15: the true pair scores higher, the noise pair lower, and the loss drops to 1.242. Word2vec is this nudge, repeated millions of times.

Figure 1 · Diagram

Reading it: the sentence becomes (center, context) pairs by a sliding window. Each true pair is a positive example; alongside it we draw a few random words as negatives ("bank" should not predict "throne"). The update pulls true pairs together and pushes noise pairs apart, then moves on to the next pair. The loop at the right is the whole training process.

In code: skipgram_pairs slides the window over a sentence and returns every (center, context) pair; build_corpus writes the lesson's synthetic sentences for it to slide over.

Chapter 2

The math: skip-gram with negative sampling (SGNS)

Every word has two vectors: vᵥ when it's the center and uᵥ when it's a context. For a center c, its true context o and k noise words n₁ … nₖ:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
L the loss for one training pair: how wrong the guesses are ≥ 0; 0 is perfect
c, o the center word and its true context word word ids
vc the center word's vector d numbers (d = 16 in this lesson)
uo the true context word's vector d numbers
nᵢ, u_{nᵢ} the i-th noise word and its vector k of them, d numbers each
k how many noise words per true pair 5 here; 5 to 20 in practice
· dot product one number
σ sigmoid, squashes a score into a probability 0 to 1
e Euler's number, ≈ 2.718
log natural logarithm; −log p is large when p is small
Σ add up over i = 1 … k

In words: the loss is small when the model gives the true pair a high probability and every noise pair a low one.

On the example: L = −log σ(0) − log σ(−0) = 0.693 + 0.693 = 1.386; after the nudge, −log σ(0.15) − log σ(0.15) = 0.621 + 0.621 = 1.242.

In Python:

import math
def sigma(x):
    # σ(x) = 1 / (1 + e^(-x))
    return 1 / (1 + math.exp(-x))
def dot(a, b):
    return sum(a_i * b_i for a_i, b_i in zip(a, b))
def L(v_c, u_o, noise):
    # -log σ(u_o · v_c)
    return (-math.log(sigma(dot(u_o, v_c)))
            # - Σ_i log σ(-u_ni · v_c)
            - sum(math.log(sigma(-dot(u_n, v_c))) for u_n in noise))
round(L((1, 0), (0, 1), [(0, -1)]), 3)  # → 1.386
# after the nudge
round(L((1, 0.1), (0.05, 1), [(-0.05, -1)]), 3)  # → 1.242

The nudge follows the gradient: for each vector, the direction in which a small change would increase the loss fastest. We step the opposite way. For this loss the gradients are short enough to derive by hand (sgns_grads), using the fact that the slope of −log σ(x) is σ(x) − 1:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Range
g_o how far the true pair's probability is from 1, as a negative number −1 to 0
gᵢ the i-th noise pair's probability, which should be 0 0 to 1
∂L/∂vc gradient: how the loss changes as each number in vc changes d numbers

In words: move the center vector toward the true context in proportion to how wrong that guess was, and away from each noise word in proportion to how much it was wrongly believed.

On the example: g_o = 0.5 − 1 = −0.5 and g₁ = 0.5, so ∂L/∂vc = −0.5·(0, 1) + 0.5·(0, −1) = (0, −1). Stepping against it with learning rate 0.1: vc = (1, 0) − 0.1·(0, −1) = (1, 0.1).

In Python:

import math
def sigma(x):
    return 1 / (1 + math.exp(-x))
v_c, u_o, u_n = (1, 0), (0, 1), (0, -1)
# σ(u_o · v_c) - 1
g_o = sigma(sum(u * v for u, v in zip(u_o, v_c))) - 1
# σ(u_n1 · v_c)
g_1 = sigma(sum(u * v for u, v in zip(u_n, v_c)))
g_o, g_1  # → (-0.5, 0.5)
# ∂L/∂v_c = g_o u_o + Σ_i g_i u_ni
grad = [g_o * o + g_1 * n for o, n in zip(u_o, u_n)]
grad  # → [0.0, -1.0]
# step against it, learning rate 0.1
[v - 0.1 * g for v, g in zip(v_c, grad)]  # → [1.0, 0.1]

Noise words are drawn in proportion to their count raised to the 3/4 power:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Range
P(w) chance that word w is drawn as a noise word 0 to 1; all add to 1
count(w) how many times w appears in the training text ≥ 1
^0.75 raise to the power 3/4, which shrinks big counts more than small ones
w′ every word in the vocabulary, in turn
Σ over w′ add up over the whole vocabulary, so the shares sum to 1

In words: common words are drawn as noise more often, but the 3/4 power gives rare words a boost.

On an example: counts 1 and 16 become 1^0.75 = 1 and 16^0.75 = 8, so the shares are 1/9 and 8/9 (0.111 and 0.889) instead of 1/17 and 16/17 (0.059 and 0.941).

In Python:

counts = [1, 16]
# count(w)^0.75
weights = [count ** 0.75 for count in counts]
weights  # → [1.0, 8.0]
# divide by Σ over w′ so the shares add to 1
[round(w / sum(weights), 3) for w in weights]  # → [0.111, 0.889]
# without the 3/4 power
[round(c / sum(counts), 3) for c in counts]  # → [0.059, 0.941]

In code: sgns_loss computes L for one center, one context and k noise words, sgns_update takes one step against the gradient, and noise_distribution builds P(w). train_sgns runs the whole loop in mini-batches and returns a WordVectors table; gradient_check confirms the hand-derived gradients against a numerical estimate.

Why it matters a full softmax over the vocabulary (primer.ml.attention explains softmax) would score every word, 100,000 or more, for every training pair. Negative sampling scores k + 1. That trick is what made training on billions of words practical in 2013, and the same pull-together, push-apart shape reappears in every modern embedding model (primer.ml.embeddings.contrastive).

Chapter 3

The famous result: vector arithmetic

Everyday picture on a street map, "walk two blocks east" is the same walk wherever you start. word2vec's map works the same way: the walk from "man" to "woman" is roughly the same walk as from "king" to "queen", because in both cases the contexts change from he/his to she/her.

Tiny example with made-up 2-D vectors: man = (1, 0), woman = (1, 1), king = (3, 0). Then king − man + woman = (3 − 1 + 1, 0 − 0 + 1) = (3, 1), exactly where queen = (3, 1) would be.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here
a, b, c the three given words, e.g. king, man, woman
v_a − v_b + v_c start at king, remove the man direction, add the woman direction
cos cosine similarity (primer.ml.embeddings.similarity)
arg max over w the word w that makes the cosine largest
w ∉ {a, b, c} skip the three input words themselves
ŵ the answer

In words: the answer is the word whose vector points most nearly the same way as king minus man plus woman, not counting those three words.

On the example: v = (3, 1) and queen = (3, 1) point the same way, so the cosine is 1, and queen wins.

In Python:

import math
def cos(a, b):
    dot = sum(a_i * b_i for a_i, b_i in zip(a, b))
    return dot / (math.sqrt(sum(a_i ** 2 for a_i in a)) * math.sqrt(sum(b_i ** 2 for b_i in b)))
v = {"man": (1, 0), "woman": (1, 1), "king": (3, 0), "queen": (3, 1)}
a, b, c = "king", "man", "woman"
# v_a - v_b + v_c
target = [x_a - x_b + x_c for x_a, x_b, x_c in zip(v[a], v[b], v[c])]
target  # → [3, 1]
# w ∉ {a, b, c}
candidates = [w for w in v if w not in {a, b, c}]
# arg max of the cosine
w_hat = max(candidates, key=lambda w: cos(v[w], target))
w_hat, round(cos(v[w_hat], target), 2)  # → ('queen', 1.0)

This lesson trains on a small synthetic corpus built from three independent attributes (male/female, royal/common, adult/child), so the effect appears in seconds on a laptop.

Figure 2 · Drawn from the lesson's code

−0.4 −0.2 0.0 0.2 0.4 principal component 1 −0.4 −0.2 0.0 0.2 0.4 principal component 2 Learned word vectors: male→female arrows are parallel king queen prince princess man woman boy girl

In a 2-D PCA view, the arrows king to queen, man to woman, prince to princess and boy to girl all point the same way with nearly equal length

Reading it: these are the learned 16-dimensional vectors, squashed to 2-D with PCA (a way to find the two directions along which the points are most spread out, and draw the points along those; primer.notation explains it). Arrows join male-female pairs: king→queen, man→woman, prince→princess, boy→girl. The arrows are roughly parallel and about the same length. That parallelism is the analogy: add the man→woman arrow to king and you land near queen. The squashing loses some structure, so treat the picture as intuition, not measurement.

In code: analogy returns the word nearest v_a − v_b + v_c, skipping the three inputs; nearest lists a word's closest neighbours by cosine, and nearest_to_vector does the same for any vector.

Why it matters this was the first clear evidence that learned vectors capture meaning as geometry. It's why "embedding" became the default way to represent text, and why vector arithmetic (averaging, subtracting, finding nearest neighbours) works on meaning.

Chapter 4

Counting instead of predicting: PPMI and GloVe

Everyday picture instead of playing the guessing game, keep a tally sheet: for every pair of words, count how often they appear near each other. Then ask which pairs turn up together more than chance would predict.

Tiny example two words that each appear only with themselves: counts [[2, 0], [0, 2]]. Out of 4 sightings, each word appears half the time (P = 0.5), and each diagonal pair also appears half the time. Chance would predict 0.5 · 0.5 = 0.25, but we see 0.5, twice as often, so the score is log(0.5 / 0.25) = log 2 = 0.693. The off-diagonal pairs never co-occur, and their score is clipped to 0.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Range
w, c a word and a context word
P(w, c) share of all co-occurrences that are this pair 0 to 1
P(w), P(c) share of co-occurrences involving w (or c) at all 0 to 1
P(w)P(c) how often the pair would appear if they were unrelated
log natural logarithm; log 1 = 0 means "exactly chance"
PMI pointwise mutual information: log of seen ÷ expected any; > 0 means "together more than chance"
max(·, 0) keep positive values, replace negatives with 0

In words: PMI asks how many times more often two words appear together than they would by chance, on a log scale; PPMI keeps only the "more often" part.

On the example: PMI = log(0.5 / (0.5 · 0.5)) = log 2 = 0.693 on the diagonal; the off-diagonal pairs have P(w, c) = 0, so PMI = −∞ and PPMI = 0.

In Python:

import math
# co-occurrence counts
X = [[2, 0], [0, 2]]
# 4 sightings
total = sum(sum(row) for row in X)
# P(w): share of each row
P_w = [sum(row) / total for row in X]
# P(c): share of each column
P_c = [sum(col) / total for col in zip(*X)]
def ppmi(w, c):
    P_wc = X[w][c] / total
    # log 0 = -∞
    pmi = math.log(P_wc / (P_w[w] * P_c[c])) if P_wc else -math.inf
    # PPMI = max(PMI, 0)
    return max(pmi, 0.0)
[[round(ppmi(w, c), 3) for c in range(2)] for w in range(2)]  # → [[0.693, 0.0], [0.0, 0.693]]

The PPMI table has one row per word, as many columns as the vocabulary, and is mostly zeros. SVD (singular value decomposition, see primer.notation) compresses it into a few columns that keep its main patterns; those few numbers per row are the word's embedding.

Figure 4 · Diagram

Reading it: the same text goes in as for word2vec, but instead of a training loop there are three bulk steps: count, rescale by chance, compress. Levy and Goldberg (2014) showed that word2vec's guessing game is secretly doing the same thing: its vectors factorize a shifted PMI table.

In code: cooccurrence counts the pairs in a window, ppmi rescales the table by chance and clips negatives to 0, and ppmi_svd_embeddings runs all three steps and keeps d columns of the SVD.

GloVe (Pennington et al., 2014) is the best-known counting method. It fits vectors so that each dot product predicts the log of the co-occurrence count:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape
i, j two words in the vocabulary
wᵢ word i's vector d numbers
w̃ⱼ (w-tilde) word j's vector when it plays the context role d numbers
bᵢ, b̃ⱼ one extra adjustable number per word (an offset for very common words) one number each
Xᵢⱼ how many times words i and j appear together a count
log natural logarithm, which tames huge counts
≈ "should be close to"; training shrinks the gap

In words: two words' vectors should have a dot product (plus offsets) that tells you how often they meet, on a log scale.

On an example: if "ice" and "cold" appear together 20 times, training nudges the vectors until w_ice · w̃_cold + b_ice + b̃_cold ≈ log 20 = 3.0. With 2-D vectors w_ice = (1, 1) and w̃_cold = (1, 0.5) and offsets 0.25 each, the left side is 1.5 + 0.25 + 0.25 = 2.0, still 1.0 short, so training keeps pushing the two vectors to agree more.

In Python:

import math
# log X_ij: the target for 20 co-occurrences
round(math.log(20), 1)  # → 3.0
w_ice, w_cold_tilde = (1, 1), (1, 0.5)
b_ice, b_cold_tilde = 0.25, 0.25
# w_i · w̃_j + b_i + b̃_j
left = sum(a * b for a, b in zip(w_ice, w_cold_tilde)) + b_ice + b_cold_tilde
# the left side, and how far it still falls short
left, round(math.log(20) - left, 1)  # → (2.0, 1.0)

Figure 3 · Drawn from the lesson's code

he his him himself she her hers herself crown throne palace castle village farm market street adult married works old young school plays toy king queen prince princess man woman boy girl PPMI: target words (rows) vs. context words (columns) 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 1.6 PPMI

A block pattern: male words light up under he and his, female under she and her, royals under crown and throne, children under young and school

Reading it: rows are the eight target words, columns are the context words, and brighter cells mean higher PPMI (they meet more than chance). You can read the attributes straight off the table: the royal words light up under "crown" and "throne", the female words under "she" and "her", the children under "young" and "school". SVD compresses exactly these block patterns into a few directions, and ppmi_svd_embeddings then solves king − man + woman = queen just like word2vec.

Why it matters counting and predicting are two routes to the same geometry. That's reassuring: embeddings aren't magic, they're compressed co-occurrence statistics.

Chapter 5

The limit: one vector per word

Everyday picture a phone book lists one address per name. If two different people are called "Bank", the book can only print one address, somewhere between the two. A static embedding table has the same problem: "bank" gets one vector for both the riverbank and the bank account.

Tiny example after training on this lesson's corpus, the single "bank" vector has cosine about 0.52 with river words and 0.50 with money words: it sits halfway between.

The fix is context. A transformer doesn't look words up once and stop. It computes a fresh vector for every word in every sentence by attending over the neighbours (primer.ml.attention): the new vector is a weighted average of the sentence's vectors,

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape / range
j which word of the sentence 1 to the sentence length
xⱼ the static vector of the j-th word d numbers
αⱼ (alpha) how much attention "bank" pays to word j (from softmax) 0 to 1
Σⱼ αⱼ = 1 the attention weights are shares that add up to 1
bank′ the new, context-aware vector for "bank" d numbers

In words: the new "bank" is a blend of the sentence's words, weighted by how relevant each one is.

On the example: in "river bank fish", "bank" blends in "river" and "fish", and its cosine with the river words climbs from 0.52 to about 0.83. The same move by hand, in 2-D where the first axis is "river" and the second is "money": river = (1, 0), bank = (1, 1), fish = (1, 0), and α = (0.25, 0.5, 0.25). Then bank′ = 0.25·(1, 0) + 0.5·(1, 1) + 0.25·(1, 0) = (1, 0.5), and its cosine with the river axis climbs from 0.71 to 0.89.

In Python:

import math
# river, bank, fish: axis 1 is "river", axis 2 "money"
x = [(1, 0), (1, 1), (1, 0)]
# attention weights, Σ_j α_j = 1
alpha = [0.25, 0.5, 0.25]
# Σ_j α_j x_j
bank_new = [sum(a_j * x_j[i] for a_j, x_j in zip(alpha, x)) for i in range(2)]
bank_new  # → [1.0, 0.5]
def cos_with_river(v):
    # cosine with (1, 0)
    return v[0] / math.sqrt(v[0] ** 2 + v[1] ** 2)
round(cos_with_river(x[1]), 2), round(cos_with_river(bank_new), 2)  # → (0.71, 0.89)

Figure 6 · Diagram

Reading it: on the left a lookup table has one row for "bank", so every sentence gets the same vector. On the right the same input goes through attention, which blends in the surrounding words, so the output depends on the sentence. contextual_bank_demo runs exactly this with the learned vectors and one attention step.

Figure 5 · Drawn from the lesson's code

static 'bank' 'bank' in river bank fish 'bank' in bank loan cash 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 cosine similarity One attention step separates the senses of 'bank' similarity to river sense similarity to money sense

Static bank is about equally close to the river and money senses (0.52 vs 0.50); in context it tilts to 0.83 river or 0.85 money

Reading it: bars show cosine similarity to the river sense (probe words water, shore, boat) and to the money sense (money, deposit, account). The static vector is about equally similar to both. After one attention step over "river bank fish" it tilts to the river sense, and over "bank loan cash" it tilts to money. Same word, different vector, decided by context.

Why it matters this is why BERT (2018) and every embedding model since are built on transformers. Search, RAG and classification all need "Python the language" and "python the snake" to be different points on the map.

Test yourself

5 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Q: What does skip-gram predict, and where do its training labels come from?Think it through, then reveal

Each center word predicts the words within a window around it. The labels come from the text itself (which words actually appeared nearby), so no human labeling is needed.

Question 2Q: Why negative sampling instead of a softmax over the vocabulary?Think it through, then reveal

A full softmax needs a score for every vocabulary word for every training pair. Negative sampling turns it into k + 1 yes/no decisions (true context vs. a few noise words), which is orders of magnitude cheaper and works as well for learning vectors.

Question 3Q: Why does king − man + woman land near queen?Think it through, then reveal

The contexts that separate king from queen (he/his vs. she/her) are the same ones that separate man from woman, so training makes those differences the same direction. Adding that direction to "king" moves it toward "queen".

Question 4Q: What's the relationship between word2vec and GloVe?Think it through, then reveal

Both encode co-occurrence statistics. GloVe explicitly fits log co-occurrence counts; skip-gram with negative sampling implicitly factorizes a shifted PMI matrix. PPMI + SVD gives similar vectors.

Question 5Q: What's the main limitation of static word embeddings and what fixed it?Think it through, then reveal

One vector per word regardless of meaning, so "bank" blends its senses. Contextual embeddings from transformers compute a vector per word in context, so each sense gets its own representation.

Primary sources

The papers behind this lesson

Mikolov, Chen, Corrado and Dean, Efficient Estimation of Word Representations in Vector Space (2013)

, with Mikolov et al., Distributed Representations of Words and Phrases and their Compositionality (2013): Introduced skip-gram and negative sampling, and the vector-arithmetic analogies.

Read the annotated companion →The paper ↗
Pennington, Socher and Manning, GloVe: Global Vectors for Word Representation (2014)

Fit word vectors directly to log co-occurrence counts, the counting route to the same geometry.

The paper ↗
Levy and Goldberg, Neural Word Embedding as Implicit Matrix Factorization (2014)

Proved that skip-gram with negative sampling implicitly factorizes a shifted PMI matrix, uniting the predicting and counting families.

The paper ↗
Devlin, Chang, Lee and Toutanova, BERT (2018)

Made contextual embeddings mainstream: one vector per word per sentence, computed by a transformer.

The paper ↗

Researcher's shelf

Further reading

  • Mikolov et al., Efficient Estimation of Word Representations in Vector Space (2013): https://arxiv.org/abs/1301.3781
  • Mikolov et al., Distributed Representations of Words and Phrases and their Compositionality (negative sampling, 2013): https://arxiv.org/abs/1310.4546
  • Goldberg and Levy, word2vec Explained (the gradient derivation): https://arxiv.org/abs/1402.3722
  • Levy and Goldberg, Neural Word Embedding as Implicit Matrix Factorization (2014): https://papers.nips.cc/paper/5477-neural-word-embedding-as-implicit-matrix-factorization
  • GloVe project page (Stanford NLP): https://nlp.stanford.edu/projects/glove/
  • Devlin et al., BERT (contextual embeddings, 2018): https://arxiv.org/abs/1810.04805
  • Jay Alammar, The Illustrated Word2vec: https://jalammar.github.io/illustrated-word2vec/

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.