The lesson in one minute
What you'll be able to explain
- Attention is permutation-equivariant: shuffle the words and the outputs shuffle the same way, so without position information "dog bites man" pools to exactly the same vector as "man bites dog".
- The original transformer adds fixed sine/cosine codes; GPT-2 and BERT learn a position table that ends at max_len.
- Modern LLMs use RoPE: rotate q and k by a position-dependent angle so the score depends only on relative distance. It also makes context extension (interpolation, NTK/YaRN) practical.
Level 1
The practitioner's guide
In one sentence
Positional encoding is how a transformer learns where each token sits, since attention on its own treats the input as an unordered bag; the scheme a model uses decides how long a context it can read, whether that length can be stretched after training, and what happens when you go past it.
When you need it
You never add positions yourself; the model's authors
chose a scheme before training and it is baked into the weights. You meet it
the day a length matters: a model card says 128K tokens and you wonder how
far to trust it, a self-hosted model produces nonsense past a certain prompt
length, an embedding model rejects documents longer than its limit, or you
want to fine-tune a model to read documents longer than it was trained on.
The tell: quality that is fine at 4,000 tokens and falls off a cliff at
some longer length, with nothing in the logs. The number behind the whole
topic, from this lesson's order_similarity: with no positions, the pooled
attention outputs for "dog bites man" and "man bites dog" have cosine
similarity exactly 1.000. The model cannot tell them apart. Nothing about
this concerns you when prompts stay well inside the length the model was
trained at.
Your options
The first four are choices the model's authors made; you choose among them by choosing a model. The last three are what you or the vendor can do about length afterwards. From the simplest to the most committed:
| Option | What it does | What it gives you | What it costs | Where it lives |
|---|---|---|---|---|
| A learned position table (GPT-2, BERT) | One learned vector per seat, added to the token's vector | Simple and effective inside the trained length | A hard ceiling: no row exists past the table's end (1,024 rows for GPT-2, 512 for BERT) | The architecture; max_position_embeddings on the card |
| Sinusoidal codes (the 2017 transformer) | A fixed sine and cosine fingerprint per position, added to the token | No parameters; a code exists for any position | Position is mixed into content; little used in new models | The architecture |
| RoPE (Llama, Mistral, Qwen) | Rotates each query and key by an angle that grows with position, so scores depend on the distance between tokens | Relative distance for free, and a context that can be stretched after training | Angles past the trained range are unfamiliar; extension needs scaling | The architecture; rope_theta and rope_parameters on the card |
| ALiBi | Subtracts a penalty proportional to distance from each score; no position vectors at all | Trained at 1,024 tokens, it extrapolates to 2,048; 11% faster and 11% less memory than sinusoidal in its paper | A built-in preference for nearby tokens; fewer models use it | The architecture |
| Inference-time scaling (linear or dynamic NTK) | Rescales positions or the RoPE base: linear scaling at every length, dynamic scaling only once a prompt exceeds the trained length | A modest stretch with no training at all | Quality drops as the stretch grows | The serving engine's config: rope_type and factor |
| Extension with a short fine-tune (PI, YaRN) | Scales positions into the trained range, then fine-tunes briefly on long text | 8× longer context: Llama to 32,768 tokens within 1,000 steps (PI); YaRN needs 10× fewer tokens than earlier methods | Long documents to train on, a training run, an evaluation at length | Your training stack |
| Staged long-context pretraining | The vendor grows the window during pretraining, checking a needle-in-a-haystack test at each stage | The genuine article: Llama 3 went from 8K to 128K in six stages | About 800B training tokens for Llama 3 405B; reaches you as a number on the card | The vendor |
How to choose
Start from the model in front of you.
- A hosted API: nothing to choose, but the window on the card is the length the vendor trained and tested, not a promise about your task. Measure at the lengths you will use.
- Picking an open model: read
max_position_embeddingsandrope_parameters. An entry with afactormeans the model was trained shorter and stretched, so test it at the lengths you will use rather than trusting the stretched number. - Serving a model beyond its trained length: do not raise the engine's context limit on its own. Set the scaling the checkpoint expects (or a dynamic scaling if none is given), and test before shipping.
- Documents longer than the model's window, and you can train: position interpolation or YaRN plus a fine-tune on long text, with a needle-in-a-haystack check and your own task at the target length.
- An encoder or embedding model with a learned table: the limit is hard.
Chunk documents to fit (
primer.agents.rag). - Whatever you pick, evaluate at the length you will run, with the answer placed at the start, the middle and the end of the prompt.
What it costs
Positions are nearly free to compute; length is what costs.
- Parameters and compute. GPT-2's table is 1,024 × 768 = 786,432 numbers, under 1% of its 124 million parameters. RoPE and ALiBi add none, and RoPE is a few element-wise multiplications per query and key, negligible next to attention itself.
- Stretching. Position interpolation for 4,096 to 32,768 tokens multiplies every position by 1/8 (32,000 becomes 4,000); NTK-aware scaling instead raises the RoPE base from 10,000 to 82,685 for a 128-wide head, leaving the fastest pair untouched and slowing the slowest 8×. Both are followed by a short fine-tune: within 1,000 steps in the PI paper. Llama 3's authors set the base to 500,000 and spent about 800B tokens taking the 405B model from 8K to 128K.
- Quality. RoPE gives nearby tokens a head start (the RoFormer companion's
long-term decay), and a long window is not used evenly: Liu et al. (Lost
in the Middle, 2023) found models use information best at the start or
the end of a long prompt and worst in the middle (
primer.agents.context).
What breaks
- Past the table. A learned-table model has no vector for position 1,024 if its table has 1,024 rows: the request fails or the library truncates. Chunk the input.
- Past the trained angles. RoPE degrades silently rather than failing: this lesson's figure shows the slowest pair's angle leaving the trained band right after 4K tokens and reaching 8× its top by 32K. The PI paper reports that plain extrapolation can produce catastrophically large attention scores. Scale, then fine-tune.
- A checkpoint served with the wrong scaling. An extended model whose
rope_parametersare dropped or mistyped reads positions it never learned. Copy the config with the weights. - Interpolating too far. Interpolation slows every frequency equally, which also blurs the fast ones that carry local word order. NTK-aware scaling and YaRN keep the fast frequencies, which is why they are preferred for large stretches.
- Pooling that forgets order. Average the token vectors of an order-blind model and "dog bites man" equals "man bites dog" exactly. Order survives only if positions are injected, or an encoding step that is not permutation-blind comes first.
- Trusting the causal mask for order. In a decoder the mask alone lets the sentences differ (cosine 0.524 in this lesson's run), but it is a weak signal, not a position encoding.
In the wild
The original transformer (Vaswani et al., 2017) used
sinusoidal codes; GPT-2 and BERT switched to learned tables, which is why
each has a fixed maximum length. Llama, Mistral and Qwen use RoPE, and Llama
3 raised its base to 500,000 with 8 key-value heads beside it. Hugging Face
model configs carry the scheme as rope_parameters with a rope_type of
linear, dynamic, yarn, longrope or llama3, a factor and an
original_max_position_embeddings; vLLM's --max-model-len sets the
served length and --hf-overrides changes that config at load time. YaRN
(Peng et al., 2023) is the extension recipe behind many long-context
checkpoints, ALiBi (Press et al., 2021) the alternative that extrapolates
without vectors, and the NoPE study (Kazemnejad et al., 2023) found that a
decoder can generalize to longer inputs with no explicit positions at all.
The papers are linked at the end of the lesson.
Go deeper
Level 2 shows the bag-of-words problem in three numbers, builds sinusoidal codes as a clock with many hands, checks RoPE's distance property by hand at positions 3 and 7 and again at 103 and 107, and stretches a 4K model to 32K by interpolation and by raising the base. If you only needed to read a model card or serve a model safely, you are done.
Level 2
How it works, from scratch
Level 2 starts with the problem itself, a bag of Scrabble tiles, and adds each family of positions in turn.
Opening
The problem: attention reads a bag, not a sentence
Everyday picture Pour the words of a sentence into a bag of Scrabble tiles. The bag for "dog bites man" and the bag for "man bites dog" hold exactly the same tiles. Anyone who only gets the bag cannot tell a news story from a routine dog attack. Attention, on its own, only ever gets the bag.
Tiny worked example Give each word a 2-number embedding: dog = (1, 0), bites = (0, 1), man = (1, 1). Average the words of each sentence, which is the simplest possible "sentence vector":
| sentence | sum of vectors | average |
|---|---|---|
| dog bites man | (1,0)+(0,1)+(1,1) = (2, 2) | (0.67, 0.67) |
| man bites dog | (1,1)+(0,1)+(1,0) = (2, 2) | (0.67, 0.67) |
Addition doesn't care about order, and neither does attention: every token compares itself with every other by dot product, and nothing in that computation says who came first.
Figure 2 · Diagram
flowchart LR
A["'dog bites man'"] --> E1[Embed each token]
B["'man bites dog'"] --> E2[Embed each token]
E1 --> S1["Set {dog, bites, man}"]
E2 --> S2["Set {man, bites, dog}"]
S1 --> ATT[Attention with no positions]
S2 --> ATT
ATT --> SAME[Same vectors, only reordered:<br/>the meaning difference is lost]
The math A permutation is a reshuffle of an ordered list. Written as a matrix (a grid of numbers), a permutation matrix is all zeros except one 1 in each row, and multiplying by it just reorders rows. Attention without positions is permutation-equivariant:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the sentence's token vectors, one row per word | rows dog (1,0), bites (0,1), man (1,1): a 3 × 2 grid | |
| a permutation matrix: reorders the rows of whatever it multiplies | swap rows 1 and 3 | |
| the same rows in the new order | man (1,1), bites (0,1), dog (1,0) | |
| one attention layer with no position information (here with = identity, so q = k = v = the word vector) | 3 × 2 in, 3 × 2 out | |
| both sides are exactly the same grid of numbers |
In words: "running attention on the shuffled sentence gives the same answer as running it on the original sentence and shuffling afterwards."
With the numbers: in "dog bites man", dog's query (1,0) scores the three keys (1,0), (0,1), (1,1) as 1, 0, 1; divided by √2 that is 0.707, 0, 0.707. Softmax gives weights 0.401, 0.198, 0.401, so dog's output is 0.401·(1,0) + 0.198·(0,1) + 0.401·(1,1) = (0.802, 0.599). In "man bites dog" dog sits in row 3, but it scores the same three keys in a different order, gets the same weights, and outputs the same (0.802, 0.599).
Level 3: in Python
import math
# q = k = v = the word vector, no positions
def Attn(X):
out = []
for q in X:
scores = [sum(a * b for a, b in zip(q, k)) / math.sqrt(2) for k in X]
exps = [math.exp(s) for s in scores]
weights = [e / sum(exps) for e in exps]
out.append(tuple(round(sum(w * v[c] for w, v in zip(weights, X)), 3) for c in range(2)))
return out
# dog, bites, man
X = [(1, 0), (0, 1), (1, 1)]
# P swaps rows 1 and 3: man, bites, dog
PX = [X[2], X[1], X[0]]
Attn(X) # → [(0.802, 0.599), (0.599, 0.802), (0.752, 0.752)]
# the same rows, shuffled the same way
Attn(PX) # → [(0.752, 0.752), (0.599, 0.802), (0.802, 0.599)]
Shuffle the input and you get the same outputs, shuffled the same way.
encode_sentence(..., scheme="none") runs a real attention layer and shows
the "dog" row is the same vector whether "dog" comes first or last.
Figure 1 · Drawn from the lesson's code
Without positions the two sentences are identical (similarity 1.0); sinusoidal codes and RoPE pull it just below 1, and a causal mask alone drops it to 0.52
In code: word_vector gives each word a fixed random embedding, and order_similarity pools each sentence's outputs and compares them, which gives the bars in the figure.
Why it matters Almost all meaning in language is carried by order: negation scope, who did what to whom, code. So position must be injected. There are three families, below.
Chapter 1
Sinusoidal codes (the original 2017 transformer)
Everyday picture A clock with many hands. The second hand moves fast and tells nearby moments apart; the hour hand moves slowly and tells morning from evening. Read all the hands together and every moment has a unique fingerprint. Sinusoidal codes give every position such a set of "hands" and add them to the token's embedding.
Tiny worked example With 4 dimensions there are two hands. The fast one turns 1 radian per position; the slow one turns 1/100 of a radian:
| position | sin(pos·1) | cos(pos·1) | sin(pos/100) | cos(pos/100) |
|---|---|---|---|---|
| 0 | 0.000 | 1.000 | 0.000 | 1.000 |
| 1 | 0.841 | 0.540 | 0.010 | 1.000 |
| 2 | 0.909 | −0.416 | 0.020 | 1.000 |
The fast pair already looks very different at positions 1 and 2; the slow pair barely moves, and it will only distinguish positions hundreds apart.
Figure 4 · Diagram
flowchart LR
T["token id"] --> E["embedding lookup<br/>(what the word means)"]
P["position 0,1,2..."] --> S["sin/cos at many frequencies<br/>(where the word is)"]
E --> ADD(("+"))
S --> ADD
ADD --> B["transformer blocks"]
The math and the code Two functions do the work. Picture a point walking around a circle of radius 1. The angle it has turned is measured in radians (a full turn is 2π ≈ 6.28 radians). Cosine of the angle is how far right the point is, and sine is how far up; both always lie between −1 and +1, and both repeat every full turn.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the token's position, counting from 0 | 1 | |
| how many numbers each position code has | 4 | |
| which (sin, cos) pair, counting from 0 up to | 0 and 1 | |
| , | the two columns that pair fills: even column gets sine, odd gets cosine | pair 1 fills columns 2 and 3 |
| pair 's speed in radians per position (its frequency) | , | |
| 10,000 raised to a fraction; makes each pair slower than the last | ||
| , | how far up / right a point is after turning by that many radians | |
| the number in row , column of the code table |
In words: "for position pos, each pair of columns stores where a hand turning at that pair's speed points after pos steps: sine in the first column, cosine in the second."
With the numbers: d = 4, pos = 1. Pair 0 turns 1 radian: sin 1 = 0.841, cos 1 = 0.540. Pair 1 turns 1/100 radian: sin 0.01 = 0.010, cos 0.01 = 1.000. So position 1's code is (0.841, 0.540, 0.010, 1.000), the second row of the table above.
Level 3: in Python
import math
d, pos = 4, 1
PE = []
# one (sin, cos) pair per i
for i in range(d // 2):
# ω_i: pair i's speed
omega_i = 1 / 10000 ** (2 * i / d)
# columns 2i and 2i+1
PE += [math.sin(pos * omega_i), math.cos(pos * omega_i)]
[f"{v:.3f}" for v in PE] # → ['0.841', '0.540', '0.010', '1.000']
sinusoidal_encoding(n, d) builds the whole (n × d) table in four lines. A
shift of k positions rotates every (sin, cos) pair by the same angle k·ω_i
wherever you start, so the dot product of two position codes depends only
on their distance.
Figure 3 · Drawn from the lesson's code
Left columns flip colour every few positions and right columns change slowly, so every row is a unique fingerprint and neighbours look alike
Why it matters Fixed codes need no training and exist for any position, and the distance property lets the model learn "look 3 tokens back" once and use it everywhere. But the code is mixed into the content vector itself, which later methods avoid.
Chapter 2
Learned absolute positions (GPT-2, BERT)
Everyday picture Numbered theatre seats. Every seat has its own brass plate, and the theatre has a fixed number of seats.
Tiny worked example GPT-2 has a table of 1,024 position vectors, each 768 numbers long: 786,432 extra parameters, learned like any other weight. Position 5 looks up row 5. Position 1,024 has no row at all.
Figure 5 · Diagram
flowchart LR
T["token id 318"] --> TE["token table<br/>(vocab × d)"]
P["position 5"] --> PE["position table<br/>(max_len × d)"]
TE --> ADD(("+"))
PE --> ADD
ADD --> B["transformer blocks"]
P2["position ≥ max_len"] -.-> X["no row: cannot be encoded"]
The math and the code
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example (GPT-2) |
|---|---|---|
| the token's row of the token table (what the word means) | 768 numbers | |
| row of the position table, learned during training (where it is) | 768 numbers; from 0 to 1,023 | |
| add the two lists number by number | ||
| the vector the first transformer block receives | 768 numbers |
In words: "the input vector is the word's meaning plus its seat number's learned vector."
With the numbers: with 2 dimensions, if = (0.5, −0.2) and = (0.1, 0.3), then = (0.6, 0.1). does not exist.
Level 3: in Python
E_dog = [0.5, -0.2]
# a position table with rows 0 to 3
P = [[0.0, 0.0], [0.0, 0.0], [0.0, 0.0], [0.1, 0.3]]
# x_3 = E_dog + P_3
[round(e + p, 2) for e, p in zip(E_dog, P[3])] # → [0.6, 0.1]
# no row 1024: it cannot be encoded
len(P) > 1024 # → False
primer.ml.transformer.TinyGPT uses exactly this.
Why it matters It's simple and works well inside the trained length, but it cannot extrapolate. That limitation pushed the field to RoPE.
Chapter 3
Rotary position embeddings, RoPE (Llama, Mistral, Qwen and most modern LLMs)
Everyday picture Two people point at a clock face, one at 1 o'clock and one at 4 o'clock. The angle between their arms is 90°. Move both of them on to 8 and 11 o'clock and the angle is still 90°. RoPE turns each token's query and key like clock hands, by an amount proportional to its position. When a query meets a key, only the angle between them (the distance between the tokens) affects the score, not where on the clock they started.
Tiny worked example Use 2 dimensions, so there is one pair turning 1 radian per position. Let q = k = (1, 0). Put the query at position 3 and the key at position 7:
- q turns by 3 rad: (cos 3, sin 3) = (−0.990, 0.141)
- k turns by 7 rad: (cos 7, sin 7) = (0.754, 0.657)
- score = −0.990·0.754 + 0.141·0.657 = −0.654 = cos(4)
Now move both 100 positions later, to 103 and 107. The score is still cos(107 − 103) = cos 4 = −0.654. Only the distance, 4, survived.
Figure 8 · Diagram
flowchart LR X[Token vector at position m] --> Q[q = x W_q] X --> K[k = x W_k] X --> V[v = x W_v] Q --> RQ["Rotate each pair of q<br/>by m·θ_i"] K --> RK["Rotate each pair of k<br/>by m·θ_i"] RQ --> S["Score = q'·k'<br/>depends only on distance"] RK --> S S --> SM[Softmax] --> W[Weighted sum of V<br/>V is NOT rotated] V --> W
The math and the code A rotation turns a point around the centre by some angle without changing its distance from the centre. Written as a 2 × 2 grid times a column of two numbers (a matrix-vector multiply: each output number is one row of the grid multiplied position-by-position with the input and added up):
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| pair of a query or key vector, before rotating | (1, 0) | |
| the same pair after rotating | (−0.990, 0.141) | |
| the token's position | 3 | |
| which pair, from 0 to | 0 (the only pair when d = 2) | |
| pair 's turning speed in radians per position; means | ||
| the total angle turned | 3 radians | |
| the 2 × 2 grid | the rotation: row 1 gives the new first number, row 2 the new second |
In words: "new first number = cos(angle) × old first − sin(angle) × old second; new second number = sin(angle) × old first + cos(angle) × old second, where the angle is position times the pair's speed."
With the numbers: x = (1, 0) at m = 3, θ = 1: x′ = (cos 3 · 1 − sin 3 · 0, sin 3 · 1 + cos 3 · 0) = (−0.990, 0.141).
Level 3: in Python
import math
def rotate(x, m, theta_i=1.0):
# the angle m θ_i
a = m * theta_i
# row 1 of the grid
return (math.cos(a) * x[0] - math.sin(a) * x[1],
# row 2 of the grid
math.sin(a) * x[0] + math.cos(a) * x[1])
[f"{v:.3f}" for v in rotate((1, 0), m=3)] # → ['-0.990', '0.141']
d, i = 2, 0
# θ_0 = 10000^(-2i/d)
10000 ** (-2 * i / d) # → 1.0
Why only the distance survives: turning q by angle a and k by angle b and then taking their dot product (multiply matching numbers, add them up; large when the vectors point the same way) gives the same answer as turning k alone by b − a. So
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| , | a query and a key vector, before rotating | (1, 0) and (1, 0) |
| , | the query's and the key's positions | 3 and 7 |
| "rotate every pair by its speed times " | turn by 3 radians | |
| the rotation for the distance only | turn by 4 radians | |
| the dot product of and (same as ) |
In words: "the score between a query at m and a key at n equals the score you'd get by leaving the query alone and turning the key by the gap n − m."
With the numbers: left side (−0.990, 0.141) · (0.754, 0.657) = −0.654; right side (1, 0) · (cos 4, sin 4) = cos 4 = −0.654. Same number, and the same again at positions 103 and 107.
Level 3: in Python
import math
# turn a pair by an angle (θ = 1, so the angle is the position)
def R(angle, x):
return (math.cos(angle) * x[0] - math.sin(angle) * x[1],
math.sin(angle) * x[0] + math.cos(angle) * x[1])
def dot(a, b): return sum(a_i * b_i for a_i, b_i in zip(a, b))
q = k = (1, 0)
# ⟨R_m q, R_n k⟩ with m = 3, n = 7
round(dot(R(3, q), R(7, k)), 3) # → -0.654
# ⟨q, R_(n-m) k⟩: only the distance
round(dot(q, R(7 - 3, k)), 3) # → -0.654
# 100 positions later, the same score
round(dot(R(103, q), R(107, k)), 3) # → -0.654
apply_rope does this for a vector or a whole sequence with three
element-wise lines, and no matrix multiply.
Figure 6 · Drawn from the lesson's code
The score curves for starting positions 0, 100 and 1,000 lie exactly on top of each other: only the distance matters
Figure 7 · Drawn from the lesson's code
Pair 0 wraps a full turn about every 6 tokens while the last pairs barely turn across the whole range
In code: rope_frequencies returns each pair's turning speed θ_i, which apply_rope multiplies by the position to get every angle.
Why it matters Relative distance is what language actually uses ("the word two back"), rotation keeps vector lengths unchanged, and because position lives in angles, you can stretch it after training. That last point is the next section.
Chapter 4
Extending context after training
Everyday picture A ruler marked from 0 to 4,000. To measure something 32,000 long you can shrink the object by 8× so it fits on the ruler (interpolation), or re-draw only the long-distance marks while keeping the fine millimetre marks (NTK-aware scaling).
Tiny worked example A model trained on 4,096 tokens now reads 32,768. Interpolation multiplies every position by 4096/32768 = 1/8, so position 32,000 is treated as 4,000: an angle the model has seen. Neighbours are now 1/8 of a position apart, which a short fine-tune teaches it to resolve.
Figure 10 · Diagram
flowchart LR
P["positions 0 ... 32k"] --> C{method}
C -->|interpolation| PI["pos × 4k/32k<br/>all hands slowed 8×"]
C -->|NTK-aware| NTK["raise the RoPE base<br/>slow hands slowed, fast hands kept"]
PI --> R["RoPE with angles inside<br/>the trained range"]
NTK --> R
R --> F["short fine-tune on long text"]
The math and the code Position interpolation (interpolate_positions):
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the real position in the long input | 32,000 | |
| the longest context seen in training | 4,096 | |
| the context we now want | 32,768 | |
| the position RoPE is actually given | 4,000 |
In words: "shrink every position by the ratio of old length to new length."
With the numbers: 32,000 × 4,096 / 32,768 = 32,000 / 8 = 4,000.
Level 3: in Python
pos, L_train, L_new = 32_000, 4_096, 32_768
# pos' = pos · L_train / L_new
pos * L_train / L_new # → 4000.0
NTK-aware scaling (ntk_scaled_base) changes the base instead:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the 10,000 inside | 10,000 | |
| how many times longer the new context is | 8 | |
| the head width | 128 | |
| raised to a power just above 1 | ||
| the new, larger base | ≈ 82,685 |
In words: "multiply the base by slightly more than the stretch factor."
With the numbers: the fastest pair keeps radian per position. The slowest pair goes from $10000^{-126/128} = 1.15 \times 10^{-4}82685^{-126/128} = 1.44 \times 10^{-5}$: exactly 8× slower.
Level 3: in Python
base, s, d = 10_000, 8, 128
# s^(d/(d-2)): just above 8
round(s ** (d / (d - 2)), 2) # → 8.27
# base' = base · s^(d/(d-2))
base_new = base * s ** (d / (d - 2))
round(base_new) # → 82685
# θ_i for the last pair, 2i = 126
slowest = lambda b: b ** (-126 / 128)
f"{slowest(base):.2e}", f"{slowest(base_new):.2e}" # → ('1.15e-04', '1.44e-05')
# exactly 8× slower
round(slowest(base) / slowest(base_new), 6) # → 8.0
YaRN refines this per frequency band and is used by many long-context models. ALiBi skips position vectors entirely and subtracts a penalty proportional to distance from each attention score.
Figure 9 · Drawn from the lesson's code
Raw positions leave the trained angle band right after 4k and reach 8 times its top by 32k; interpolation and NTK scaling stay inside it
Why it matters It's how models trained on a few thousand tokens are cheaply stretched to 100k or more, instead of being pretrained again from scratch.
Opening
A subtlety: the causal mask leaks order
In a decoder, token i sees exactly i+1 tokens, so even with no position encoding ("NoPE") the model can partly infer where it is. Bidirectional attention (BERT-style) has no such crutch and is fully order-blind.
In code: encode_sentence with its causal flag set runs the same layer under primer.ml.attention.causal_mask; its pooled similarity is the "causal mask only" bar in the first figure.
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Why does a transformer need positional information at all?Think it through, then reveal
Attention compares tokens by content only, so it's permutation-equivariant: shuffling the input just shuffles the output. Without positions, word order is invisible.
Question 2How does RoPE encode position, and what property makes it attractive?Think it through, then reveal
It rotates each pair of dimensions in q and k by an angle proportional to position. The q·k score then depends only on the relative distance, and rotation preserves vector length.
Question 3Compute a RoPE score by hand: 2 dimensions, θ = 1, q = k = (1, 0), query at 3, key at 7.Think it through, then reveal
cos(7 − 3) = cos 4 ≈ −0.654, and the same at positions 103 and 107.
Question 4Why do sinusoidal codes use many frequencies?Think it through, then reveal
Fast frequencies distinguish neighbours; slow ones distinguish distant positions. Together every position gets a unique, smoothly varying code, and a shift by k is the same rotation everywhere.
Question 5What goes wrong beyond the trained context length, and how do you fix it?Think it through, then reveal
The model meets position angles or indices it never saw, and attention degrades. Position interpolation, NTK-aware scaling or YaRN rescale positions or frequencies into the trained range, usually followed by a short fine-tune.
Question 6Learned absolute position embeddings: pro and con?Think it through, then reveal
They are simple and effective within the trained length, but they cannot represent any position beyond the table size.
Primary sources
The papers behind this lesson
Introduced the transformer and, in section 3.5, the sinusoidal position codes built here.
Read on rumblr →The paper ↗Introduced RoPE: rotate queries and keys so attention scores depend only on relative distance.
Read the annotated companion →The paper ↗Showed that scaling positions down plus a short fine-tune extends RoPE models to much longer contexts.
The paper ↗Refined RoPE context extension by treating fast and slow frequencies differently.
The paper ↗Replaced position vectors with a distance penalty on attention scores that extrapolates to longer inputs.
The paper ↗Researcher's shelf
Further reading
- Vaswani et al., Attention Is All You Need (sinusoidal codes, section 3.5): https://arxiv.org/abs/1706.03762
- Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding (2021): https://arxiv.org/abs/2104.09864
- EleutherAI, Rotary Embeddings: A Relative Revolution: https://blog.eleuther.ai/rotary-embeddings/
- Chen et al., Extending Context Window of Large Language Models via Positional Interpolation (2023): https://arxiv.org/abs/2306.15595
- Peng et al., YaRN: Efficient Context Window Extension of Large Language Models (2023): https://arxiv.org/abs/2309.00071
- Press et al., Train Short, Test Long: Attention with Linear Biases (ALiBi) (2021): https://arxiv.org/abs/2108.12409
- Kazemnejad et al., The Impact of Positional Encoding on Length Generalization in Transformers (NoPE, 2023): https://arxiv.org/abs/2305.19466
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.