At a glance
Key takeaways
- Attention is permutation-equivariant: shuffle the words and the outputs shuffle the same way, so without position information "dog bites man" pools to exactly the same vector as "man bites dog".
- The original transformer adds fixed sine/cosine codes; GPT-2 and BERT learn a position table that ends at max_len.
- Modern LLMs use RoPE: rotate q and k by a position-dependent angle so the score depends only on relative distance. It also makes context extension (interpolation, NTK/YaRN) practical.
Level 2
How it works, from scratch
Level 2 starts with the problem itself, a bag of Scrabble tiles, and adds each family of positions in turn.
Opening
The problem: attention reads a bag, not a sentence
Everyday picture Pour the words of a sentence into a bag of Scrabble tiles. The bag for "dog bites man" and the bag for "man bites dog" hold exactly the same tiles. Anyone who only gets the bag cannot tell a news story from a routine dog attack. Attention, on its own, only ever gets the bag.
Tiny worked example Give each word a 2-number embedding: dog = (1, 0), bites = (0, 1), man = (1, 1). Average the words of each sentence, which is the simplest possible "sentence vector":
| sentence | sum of vectors | average |
|---|---|---|
| dog bites man | (1,0)+(0,1)+(1,1) = (2, 2) | (0.67, 0.67) |
| man bites dog | (1,1)+(0,1)+(1,0) = (2, 2) | (0.67, 0.67) |
Addition doesn't care about order, and neither does attention: every token compares itself with every other by dot product, and nothing in that computation says who came first.
Figure 1 · Diagram
flowchart LR
A["'dog bites man'"] --> E1[Embed each token]
B["'man bites dog'"] --> E2[Embed each token]
E1 --> S1["Set {dog, bites, man}"]
E2 --> S2["Set {man, bites, dog}"]
S1 --> ATT[Attention with no positions]
S2 --> ATT
ATT --> SAME[Same vectors, only reordered:<br/>the meaning difference is lost]
The math A permutation is a reshuffle of an ordered list. Written as a matrix (a grid of numbers), a permutation matrix is all zeros except one 1 in each row, and multiplying by it just reorders rows. Attention without positions is permutation-equivariant:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the sentence's token vectors, one row per word | rows dog (1,0), bites (0,1), man (1,1): a 3 × 2 grid | |
| a permutation matrix: reorders the rows of whatever it multiplies | swap rows 1 and 3 | |
| the same rows in the new order | man (1,1), bites (0,1), dog (1,0) | |
| one attention layer with no position information (here with = identity, so q = k = v = the word vector) | 3 × 2 in, 3 × 2 out | |
| both sides are exactly the same grid of numbers |
In words: "running attention on the shuffled sentence gives the same answer as running it on the original sentence and shuffling afterwards."
With the numbers: in "dog bites man", dog's query (1,0) scores the three keys (1,0), (0,1), (1,1) as 1, 0, 1; divided by √2 that is 0.707, 0, 0.707. Softmax gives weights 0.401, 0.198, 0.401, so dog's output is 0.401·(1,0) + 0.198·(0,1) + 0.401·(1,1) = (0.802, 0.599). In "man bites dog" dog sits in row 3, but it scores the same three keys in a different order, gets the same weights, and outputs the same (0.802, 0.599).
Level 3: in Python
import math
# q = k = v = the word vector, no positions
def Attn(X):
out = []
for q in X:
scores = [sum(a * b for a, b in zip(q, k)) / math.sqrt(2) for k in X]
exps = [math.exp(s) for s in scores]
weights = [e / sum(exps) for e in exps]
out.append(tuple(round(sum(w * v[c] for w, v in zip(weights, X)), 3) for c in range(2)))
return out
# dog, bites, man
X = [(1, 0), (0, 1), (1, 1)]
# P swaps rows 1 and 3: man, bites, dog
PX = [X[2], X[1], X[0]]
Attn(X) # → [(0.802, 0.599), (0.599, 0.802), (0.752, 0.752)]
# the same rows, shuffled the same way
Attn(PX) # → [(0.752, 0.752), (0.599, 0.802), (0.802, 0.599)]
Shuffle the input and you get the same outputs, shuffled the same way.
encode_sentence(..., scheme="none") runs a real attention layer and shows
the "dog" row is the same vector whether "dog" comes first or last.
Figure 2 · Drawn from the lesson's code
Without positions the two sentences are identical (similarity 1.0); sinusoidal codes and RoPE pull it just below 1, and a causal mask alone drops it to 0.52
In code: word_vector gives each word a fixed random embedding, and order_similarity pools each sentence's outputs and compares them, which gives the bars in the figure.
Why it matters Almost all meaning in language is carried by order: negation scope, who did what to whom, code. So position must be injected. There are three families, below.
Chapter 1
Sinusoidal codes (the original 2017 transformer)
Everyday picture A clock with many hands. The second hand moves fast and tells nearby moments apart; the hour hand moves slowly and tells morning from evening. Read all the hands together and every moment has a unique fingerprint. Sinusoidal codes give every position such a set of "hands" and add them to the token's embedding.
Tiny worked example With 4 dimensions there are two hands. The fast one turns 1 radian per position; the slow one turns 1/100 of a radian:
| position | sin(pos·1) | cos(pos·1) | sin(pos/100) | cos(pos/100) |
|---|---|---|---|---|
| 0 | 0.000 | 1.000 | 0.000 | 1.000 |
| 1 | 0.841 | 0.540 | 0.010 | 1.000 |
| 2 | 0.909 | −0.416 | 0.020 | 1.000 |
The fast pair already looks very different at positions 1 and 2; the slow pair barely moves, and it will only distinguish positions hundreds apart.
Figure 3 · Diagram
flowchart LR
T["token id"] --> E["embedding lookup<br/>(what the word means)"]
P["position 0,1,2..."] --> S["sin/cos at many frequencies<br/>(where the word is)"]
E --> ADD(("+"))
S --> ADD
ADD --> B["transformer blocks"]
The math and the code Two functions do the work. Picture a point walking around a circle of radius 1. The angle it has turned is measured in radians (a full turn is 2π ≈ 6.28 radians). Cosine of the angle is how far right the point is, and sine is how far up; both always lie between −1 and +1, and both repeat every full turn.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the token's position, counting from 0 | 1 | |
| how many numbers each position code has | 4 | |
| which (sin, cos) pair, counting from 0 up to | 0 and 1 | |
| , | the two columns that pair fills: even column gets sine, odd gets cosine | pair 1 fills columns 2 and 3 |
| pair 's speed in radians per position (its frequency) | , | |
| 10,000 raised to a fraction; makes each pair slower than the last | ||
| , | how far up / right a point is after turning by that many radians | |
| the number in row , column of the code table |
In words: "for position pos, each pair of columns stores where a hand turning at that pair's speed points after pos steps: sine in the first column, cosine in the second."
With the numbers: d = 4, pos = 1. Pair 0 turns 1 radian: sin 1 = 0.841, cos 1 = 0.540. Pair 1 turns 1/100 radian: sin 0.01 = 0.010, cos 0.01 = 1.000. So position 1's code is (0.841, 0.540, 0.010, 1.000), the second row of the table above.
Level 3: in Python
import math
d, pos = 4, 1
PE = []
# one (sin, cos) pair per i
for i in range(d // 2):
# ω_i: pair i's speed
omega_i = 1 / 10000 ** (2 * i / d)
# columns 2i and 2i+1
PE += [math.sin(pos * omega_i), math.cos(pos * omega_i)]
[f"{v:.3f}" for v in PE] # → ['0.841', '0.540', '0.010', '1.000']
sinusoidal_encoding(n, d) builds the whole (n × d) table in four lines. A
shift of k positions rotates every (sin, cos) pair by the same angle k·ω_i
wherever you start, so the dot product of two position codes depends only
on their distance.
Figure 4 · Drawn from the lesson's code
Left columns flip colour every few positions and right columns change slowly, so every row is a unique fingerprint and neighbours look alike
Why it matters Fixed codes need no training and exist for any position, and the distance property lets the model learn "look 3 tokens back" once and use it everywhere. But the code is mixed into the content vector itself, which later methods avoid.
Chapter 2
Learned absolute positions (GPT-2, BERT)
Everyday picture Numbered theatre seats. Every seat has its own brass plate, and the theatre has a fixed number of seats.
Tiny worked example GPT-2 has a table of 1,024 position vectors, each 768 numbers long: 786,432 extra parameters, learned like any other weight. Position 5 looks up row 5. Position 1,024 has no row at all.
Figure 5 · Diagram
flowchart LR
T["token id 318"] --> TE["token table<br/>(vocab × d)"]
P["position 5"] --> PE["position table<br/>(max_len × d)"]
TE --> ADD(("+"))
PE --> ADD
ADD --> B["transformer blocks"]
P2["position ≥ max_len"] -.-> X["no row: cannot be encoded"]
The math and the code
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example (GPT-2) |
|---|---|---|
| the token's row of the token table (what the word means) | 768 numbers | |
| row of the position table, learned during training (where it is) | 768 numbers; from 0 to 1,023 | |
| add the two lists number by number | ||
| the vector the first transformer block receives | 768 numbers |
In words: "the input vector is the word's meaning plus its seat number's learned vector."
With the numbers: with 2 dimensions, if = (0.5, −0.2) and = (0.1, 0.3), then = (0.6, 0.1). does not exist.
Level 3: in Python
E_dog = [0.5, -0.2]
# a position table with rows 0 to 3
P = [[0.0, 0.0], [0.0, 0.0], [0.0, 0.0], [0.1, 0.3]]
# x_3 = E_dog + P_3
[round(e + p, 2) for e, p in zip(E_dog, P[3])] # → [0.6, 0.1]
# no row 1024: it cannot be encoded
len(P) > 1024 # → False
primer.ml.transformer.TinyGPT uses exactly this.
Why it matters It's simple and works well inside the trained length, but it cannot extrapolate. That limitation pushed the field to RoPE.
Chapter 3
Rotary position embeddings, RoPE (Llama, Mistral, Qwen and most modern LLMs)
Everyday picture Two people point at a clock face, one at 1 o'clock and one at 4 o'clock. The angle between their arms is 90°. Move both of them on to 8 and 11 o'clock and the angle is still 90°. RoPE turns each token's query and key like clock hands, by an amount proportional to its position. When a query meets a key, only the angle between them (the distance between the tokens) affects the score, not where on the clock they started.
Tiny worked example Use 2 dimensions, so there is one pair turning 1 radian per position. Let q = k = (1, 0). Put the query at position 3 and the key at position 7:
- q turns by 3 rad: (cos 3, sin 3) = (−0.990, 0.141)
- k turns by 7 rad: (cos 7, sin 7) = (0.754, 0.657)
- score = −0.990·0.754 + 0.141·0.657 = −0.654 = cos(4)
Now move both 100 positions later, to 103 and 107. The score is still cos(107 − 103) = cos 4 = −0.654. Only the distance, 4, survived.
Figure 6 · Diagram
flowchart LR X[Token vector at position m] --> Q[q = x W_q] X --> K[k = x W_k] X --> V[v = x W_v] Q --> RQ["Rotate each pair of q<br/>by m·θ_i"] K --> RK["Rotate each pair of k<br/>by m·θ_i"] RQ --> S["Score = q'·k'<br/>depends only on distance"] RK --> S S --> SM[Softmax] --> W[Weighted sum of V<br/>V is NOT rotated] V --> W
The math and the code A rotation turns a point around the centre by some angle without changing its distance from the centre. Written as a 2 × 2 grid times a column of two numbers (a matrix-vector multiply: each output number is one row of the grid multiplied position-by-position with the input and added up):
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| pair of a query or key vector, before rotating | (1, 0) | |
| the same pair after rotating | (−0.990, 0.141) | |
| the token's position | 3 | |
| which pair, from 0 to | 0 (the only pair when d = 2) | |
| pair 's turning speed in radians per position; means | ||
| the total angle turned | 3 radians | |
| the 2 × 2 grid | the rotation: row 1 gives the new first number, row 2 the new second |
In words: "new first number = cos(angle) × old first − sin(angle) × old second; new second number = sin(angle) × old first + cos(angle) × old second, where the angle is position times the pair's speed."
With the numbers: x = (1, 0) at m = 3, θ = 1: x′ = (cos 3 · 1 − sin 3 · 0, sin 3 · 1 + cos 3 · 0) = (−0.990, 0.141).
Level 3: in Python
import math
def rotate(x, m, theta_i=1.0):
# the angle m θ_i
a = m * theta_i
# row 1 of the grid
return (math.cos(a) * x[0] - math.sin(a) * x[1],
# row 2 of the grid
math.sin(a) * x[0] + math.cos(a) * x[1])
[f"{v:.3f}" for v in rotate((1, 0), m=3)] # → ['-0.990', '0.141']
d, i = 2, 0
# θ_0 = 10000^(-2i/d)
10000 ** (-2 * i / d) # → 1.0
Why only the distance survives: turning q by angle a and k by angle b and then taking their dot product (multiply matching numbers, add them up; large when the vectors point the same way) gives the same answer as turning k alone by b − a. So
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| , | a query and a key vector, before rotating | (1, 0) and (1, 0) |
| , | the query's and the key's positions | 3 and 7 |
| "rotate every pair by its speed times " | turn by 3 radians | |
| the rotation for the distance only | turn by 4 radians | |
| the dot product of and (same as ) |
In words: "the score between a query at m and a key at n equals the score you'd get by leaving the query alone and turning the key by the gap n − m."
With the numbers: left side (−0.990, 0.141) · (0.754, 0.657) = −0.654; right side (1, 0) · (cos 4, sin 4) = cos 4 = −0.654. Same number, and the same again at positions 103 and 107.
Level 3: in Python
import math
# turn a pair by an angle (θ = 1, so the angle is the position)
def R(angle, x):
return (math.cos(angle) * x[0] - math.sin(angle) * x[1],
math.sin(angle) * x[0] + math.cos(angle) * x[1])
def dot(a, b): return sum(a_i * b_i for a_i, b_i in zip(a, b))
q = k = (1, 0)
# ⟨R_m q, R_n k⟩ with m = 3, n = 7
round(dot(R(3, q), R(7, k)), 3) # → -0.654
# ⟨q, R_(n-m) k⟩: only the distance
round(dot(q, R(7 - 3, k)), 3) # → -0.654
# 100 positions later, the same score
round(dot(R(103, q), R(107, k)), 3) # → -0.654
apply_rope does this for a vector or a whole sequence with three
element-wise lines, and no matrix multiply.
Figure 7 · Drawn from the lesson's code
The score curves for starting positions 0, 100 and 1,000 lie exactly on top of each other: only the distance matters
Figure 8 · Drawn from the lesson's code
Pair 0 wraps a full turn about every 6 tokens while the last pairs barely turn across the whole range
In code: rope_frequencies returns each pair's turning speed θ_i, which apply_rope multiplies by the position to get every angle.
Why it matters Relative distance is what language actually uses ("the word two back"), rotation keeps vector lengths unchanged, and because position lives in angles, you can stretch it after training. That last point is the next section.
Chapter 4
Extending context after training
Everyday picture A ruler marked from 0 to 4,000. To measure something 32,000 long you can shrink the object by 8× so it fits on the ruler (interpolation), or re-draw only the long-distance marks while keeping the fine millimetre marks (NTK-aware scaling).
Tiny worked example A model trained on 4,096 tokens now reads 32,768. Interpolation multiplies every position by 4096/32768 = 1/8, so position 32,000 is treated as 4,000: an angle the model has seen. Neighbours are now 1/8 of a position apart, which a short fine-tune teaches it to resolve.
Figure 9 · Diagram
flowchart LR
P["positions 0 ... 32k"] --> C{method}
C -->|interpolation| PI["pos × 4k/32k<br/>all hands slowed 8×"]
C -->|NTK-aware| NTK["raise the RoPE base<br/>slow hands slowed, fast hands kept"]
PI --> R["RoPE with angles inside<br/>the trained range"]
NTK --> R
R --> F["short fine-tune on long text"]
The math and the code Position interpolation (interpolate_positions):
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the real position in the long input | 32,000 | |
| the longest context seen in training | 4,096 | |
| the context we now want | 32,768 | |
| the position RoPE is actually given | 4,000 |
In words: "shrink every position by the ratio of old length to new length."
With the numbers: 32,000 × 4,096 / 32,768 = 32,000 / 8 = 4,000.
Level 3: in Python
pos, L_train, L_new = 32_000, 4_096, 32_768
# pos' = pos · L_train / L_new
pos * L_train / L_new # → 4000.0
NTK-aware scaling (ntk_scaled_base) changes the base instead:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the 10,000 inside | 10,000 | |
| how many times longer the new context is | 8 | |
| the head width | 128 | |
| raised to a power just above 1 | ||
| the new, larger base | ≈ 82,685 |
In words: "multiply the base by slightly more than the stretch factor."
With the numbers: the fastest pair keeps radian per position. The slowest pair goes from $10000^{-126/128} = 1.15 \times 10^{-4}82685^{-126/128} = 1.44 \times 10^{-5}$: exactly 8× slower.
Level 3: in Python
base, s, d = 10_000, 8, 128
# s^(d/(d-2)): just above 8
round(s ** (d / (d - 2)), 2) # → 8.27
# base' = base · s^(d/(d-2))
base_new = base * s ** (d / (d - 2))
round(base_new) # → 82685
# θ_i for the last pair, 2i = 126
slowest = lambda b: b ** (-126 / 128)
f"{slowest(base):.2e}", f"{slowest(base_new):.2e}" # → ('1.15e-04', '1.44e-05')
# exactly 8× slower
round(slowest(base) / slowest(base_new), 6) # → 8.0
YaRN refines this per frequency band and is used by many long-context models. ALiBi skips position vectors entirely and subtracts a penalty proportional to distance from each attention score.
Figure 10 · Drawn from the lesson's code
Raw positions leave the trained angle band right after 4k and reach 8 times its top by 32k; interpolation and NTK scaling stay inside it
Why it matters It's how models trained on a few thousand tokens are cheaply stretched to 100k or more, instead of being pretrained again from scratch.
Opening
A subtlety: the causal mask leaks order
In a decoder, token i sees exactly i+1 tokens, so even with no position encoding ("NoPE") the model can partly infer where it is. Bidirectional attention (BERT-style) has no such crutch and is fully order-blind.
In code: encode_sentence with its causal flag set runs the same layer under primer.ml.attention.causal_mask; its pooled similarity is the "causal mask only" bar in the first figure.
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Why does a transformer need positional information at all?Think it through, then reveal
Attention compares tokens by content only, so it's permutation-equivariant: shuffling the input just shuffles the output. Without positions, word order is invisible.
Question 2How does RoPE encode position, and what property makes it attractive?Think it through, then reveal
It rotates each pair of dimensions in q and k by an angle proportional to position. The q·k score then depends only on the relative distance, and rotation preserves vector length.
Question 3Compute a RoPE score by hand: 2 dimensions, θ = 1, q = k = (1, 0), query at 3, key at 7.Think it through, then reveal
cos(7 − 3) = cos 4 ≈ −0.654, and the same at positions 103 and 107.
Question 4Why do sinusoidal codes use many frequencies?Think it through, then reveal
Fast frequencies distinguish neighbours; slow ones distinguish distant positions. Together every position gets a unique, smoothly varying code, and a shift by k is the same rotation everywhere.
Question 5What goes wrong beyond the trained context length, and how do you fix it?Think it through, then reveal
The model meets position angles or indices it never saw, and attention degrades. Position interpolation, NTK-aware scaling or YaRN rescale positions or frequencies into the trained range, usually followed by a short fine-tune.
Question 6Learned absolute position embeddings: pro and con?Think it through, then reveal
They are simple and effective within the trained length, but they cannot represent any position beyond the table size.
Primary sources
The papers behind this lesson
Introduced the transformer and, in section 3.5, the sinusoidal position codes built here.
Read on rumblr →The paper ↗Introduced RoPE: rotate queries and keys so attention scores depend only on relative distance.
Read the annotated companion →The paper ↗Showed that scaling positions down plus a short fine-tune extends RoPE models to much longer contexts.
The paper ↗Refined RoPE context extension by treating fast and slow frequencies differently.
The paper ↗Replaced position vectors with a distance penalty on attention scores that extrapolates to longer inputs.
The paper ↗Researcher's shelf
Further reading
- Vaswani et al., Attention Is All You Need (sinusoidal codes, section 3.5): https://arxiv.org/abs/1706.03762
- Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding (2021): https://arxiv.org/abs/2104.09864
- EleutherAI, Rotary Embeddings: A Relative Revolution: https://blog.eleuther.ai/rotary-embeddings/
- Chen et al., Extending Context Window of Large Language Models via Positional Interpolation (2023): https://arxiv.org/abs/2306.15595
- Peng et al., YaRN: Efficient Context Window Extension of Large Language Models (2023): https://arxiv.org/abs/2309.00071
- Press et al., Train Short, Test Long: Attention with Linear Biases (ALiBi) (2021): https://arxiv.org/abs/2108.12409
- Kazemnejad et al., The Impact of Positional Encoding on Length Generalization in Transformers (NoPE, 2023): https://arxiv.org/abs/2305.19466
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.