rumblr Work in progressWIP

● The AI Primer · Lesson 6 · Part 1: how the model works inside

Positional information

how a transformer knows word order

You'll be able to explain Why order must be added, sinusoids and RoPE

Members · open during launch 31 min10 figures and diagrams7 interactive
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Attention is permutation-equivariant: shuffle the words and the outputs shuffle the same way, so without position information "dog bites man" pools to exactly the same vector as "man bites dog".
  2. The original transformer adds fixed sine/cosine codes; GPT-2 and BERT learn a position table that ends at max_len.
  3. Modern LLMs use RoPE: rotate q and k by a position-dependent angle so the score depends only on relative distance. It also makes context extension (interpolation, NTK/YaRN) practical.

Level 1

The practitioner's guide

Start from where you stand: the map picks the option, and the guide below explains every one.

In one sentence

Positional encoding is how a transformer learns where each token sits, since attention on its own treats the input as an unordered bag; the scheme a model uses decides how long a context it can read, whether that length can be stretched after training, and what happens when you go past it.

When you need it

You never add positions yourself; the model's authors chose a scheme before training and it is baked into the weights. You meet it the day a length matters: a model card says 128K tokens and you wonder how far to trust it, a self-hosted model produces nonsense past a certain prompt length, an embedding model rejects documents longer than its limit, or you want to fine-tune a model to read documents longer than it was trained on. The tell: quality that is fine at 4,000 tokens and falls off a cliff at some longer length, with nothing in the logs. The number behind the whole topic, from this lesson's order_similarity: with no positions, the pooled attention outputs for "dog bites man" and "man bites dog" have cosine similarity exactly 1.000. The model cannot tell them apart. Nothing about this concerns you when prompts stay well inside the length the model was trained at.

Your options

The first four are choices the model's authors made; you choose among them by choosing a model. The last three are what you or the vendor can do about length afterwards. From the simplest to the most committed:

Option What it does What it gives you What it costs Where it lives
A learned position table (GPT-2, BERT) One learned vector per seat, added to the token's vector Simple and effective inside the trained length A hard ceiling: no row exists past the table's end (1,024 rows for GPT-2, 512 for BERT) The architecture; max_position_embeddings on the card
Sinusoidal codes (the 2017 transformer) A fixed sine and cosine fingerprint per position, added to the token No parameters; a code exists for any position Position is mixed into content; little used in new models The architecture
RoPE (Llama, Mistral, Qwen) Rotates each query and key by an angle that grows with position, so scores depend on the distance between tokens Relative distance for free, and a context that can be stretched after training Angles past the trained range are unfamiliar; extension needs scaling The architecture; rope_theta and rope_parameters on the card
ALiBi Subtracts a penalty proportional to distance from each score; no position vectors at all Trained at 1,024 tokens, it extrapolates to 2,048; 11% faster and 11% less memory than sinusoidal in its paper A built-in preference for nearby tokens; fewer models use it The architecture
Inference-time scaling (linear or dynamic NTK) Rescales positions or the RoPE base: linear scaling at every length, dynamic scaling only once a prompt exceeds the trained length A modest stretch with no training at all Quality drops as the stretch grows The serving engine's config: rope_type and factor
Extension with a short fine-tune (PI, YaRN) Scales positions into the trained range, then fine-tunes briefly on long text 8× longer context: Llama to 32,768 tokens within 1,000 steps (PI); YaRN needs 10× fewer tokens than earlier methods Long documents to train on, a training run, an evaluation at length Your training stack
Staged long-context pretraining The vendor grows the window during pretraining, checking a needle-in-a-haystack test at each stage The genuine article: Llama 3 went from 8K to 128K in six stages About 800B training tokens for Llama 3 405B; reaches you as a number on the card The vendor

How to choose

Start from the model in front of you.

  • A hosted API: nothing to choose, but the window on the card is the length the vendor trained and tested, not a promise about your task. Measure at the lengths you will use.
  • Picking an open model: read max_position_embeddings and rope_parameters. An entry with a factor means the model was trained shorter and stretched, so test it at the lengths you will use rather than trusting the stretched number.
  • Serving a model beyond its trained length: do not raise the engine's context limit on its own. Set the scaling the checkpoint expects (or a dynamic scaling if none is given), and test before shipping.
  • Documents longer than the model's window, and you can train: position interpolation or YaRN plus a fine-tune on long text, with a needle-in-a-haystack check and your own task at the target length.
  • An encoder or embedding model with a learned table: the limit is hard. Chunk documents to fit (primer.agents.rag).
  • Whatever you pick, evaluate at the length you will run, with the answer placed at the start, the middle and the end of the prompt.

What it costs

Positions are nearly free to compute; length is what costs.

  • Parameters and compute. GPT-2's table is 1,024 × 768 = 786,432 numbers, under 1% of its 124 million parameters. RoPE and ALiBi add none, and RoPE is a few element-wise multiplications per query and key, negligible next to attention itself.
  • Stretching. Position interpolation for 4,096 to 32,768 tokens multiplies every position by 1/8 (32,000 becomes 4,000); NTK-aware scaling instead raises the RoPE base from 10,000 to 82,685 for a 128-wide head, leaving the fastest pair untouched and slowing the slowest 8×. Both are followed by a short fine-tune: within 1,000 steps in the PI paper. Llama 3's authors set the base to 500,000 and spent about 800B tokens taking the 405B model from 8K to 128K.
  • Quality. RoPE gives nearby tokens a head start (the RoFormer companion's long-term decay), and a long window is not used evenly: Liu et al. (Lost in the Middle, 2023) found models use information best at the start or the end of a long prompt and worst in the middle (primer.agents.context).

What breaks

  • Past the table. A learned-table model has no vector for position 1,024 if its table has 1,024 rows: the request fails or the library truncates. Chunk the input.
  • Past the trained angles. RoPE degrades silently rather than failing: this lesson's figure shows the slowest pair's angle leaving the trained band right after 4K tokens and reaching 8× its top by 32K. The PI paper reports that plain extrapolation can produce catastrophically large attention scores. Scale, then fine-tune.
  • A checkpoint served with the wrong scaling. An extended model whose rope_parameters are dropped or mistyped reads positions it never learned. Copy the config with the weights.
  • Interpolating too far. Interpolation slows every frequency equally, which also blurs the fast ones that carry local word order. NTK-aware scaling and YaRN keep the fast frequencies, which is why they are preferred for large stretches.
  • Pooling that forgets order. Average the token vectors of an order-blind model and "dog bites man" equals "man bites dog" exactly. Order survives only if positions are injected, or an encoding step that is not permutation-blind comes first.
  • Trusting the causal mask for order. In a decoder the mask alone lets the sentences differ (cosine 0.524 in this lesson's run), but it is a weak signal, not a position encoding.

In the wild

The original transformer (Vaswani et al., 2017) used sinusoidal codes; GPT-2 and BERT switched to learned tables, which is why each has a fixed maximum length. Llama, Mistral and Qwen use RoPE, and Llama 3 raised its base to 500,000 with 8 key-value heads beside it. Hugging Face model configs carry the scheme as rope_parameters with a rope_type of linear, dynamic, yarn, longrope or llama3, a factor and an original_max_position_embeddings; vLLM's --max-model-len sets the served length and --hf-overrides changes that config at load time. YaRN (Peng et al., 2023) is the extension recipe behind many long-context checkpoints, ALiBi (Press et al., 2021) the alternative that extrapolates without vectors, and the NoPE study (Kazemnejad et al., 2023) found that a decoder can generalize to longer inputs with no explicit positions at all. The papers are linked at the end of the lesson.

Go deeper

Level 2 shows the bag-of-words problem in three numbers, builds sinusoidal codes as a clock with many hands, checks RoPE's distance property by hand at positions 3 and 7 and again at 103 and 107, and stretches a 4K model to 32K by interpolation and by raising the base. If you only needed to read a model card or serve a model safely, you are done.

Level 2

How it works, from scratch

Level 2 starts with the problem itself, a bag of Scrabble tiles, and adds each family of positions in turn.

Opening

The problem: attention reads a bag, not a sentence

Shuffle the sentence yourself first; the lesson then shows why nothing moved.

Everyday picture Pour the words of a sentence into a bag of Scrabble tiles. The bag for "dog bites man" and the bag for "man bites dog" hold exactly the same tiles. Anyone who only gets the bag cannot tell a news story from a routine dog attack. Attention, on its own, only ever gets the bag.

Tiny worked example Give each word a 2-number embedding: dog = (1, 0), bites = (0, 1), man = (1, 1). Average the words of each sentence, which is the simplest possible "sentence vector":

sentence sum of vectors average
dog bites man (1,0)+(0,1)+(1,1) = (2, 2) (0.67, 0.67)
man bites dog (1,1)+(0,1)+(1,0) = (2, 2) (0.67, 0.67)

Addition doesn't care about order, and neither does attention: every token compares itself with every other by dot product, and nothing in that computation says who came first.

Figure 2 · Diagram

Reading it: follow both sentences left to right. Embedding looks up each token independently, so both sentences produce the same three vectors. Attention only compares vectors with each other; it has no notion of slot 1 versus slot 3. It sees the same set and produces the same output vectors, listed in a different order. Pool them into one sentence vector and the two sentences are identical.

The math A permutation is a reshuffle of an ordered list. Written as a matrix (a grid of numbers), a permutation matrix is all zeros except one 1 in each row, and multiplying by it just reorders rows. Attention without positions is permutation-equivariant:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the sentence's token vectors, one row per word rows dog (1,0), bites (0,1), man (1,1): a 3 × 2 grid
a permutation matrix: reorders the rows of whatever it multiplies swap rows 1 and 3
the same rows in the new order man (1,1), bites (0,1), dog (1,0)
one attention layer with no position information (here with = identity, so q = k = v = the word vector) 3 × 2 in, 3 × 2 out
both sides are exactly the same grid of numbers

In words: "running attention on the shuffled sentence gives the same answer as running it on the original sentence and shuffling afterwards."

With the numbers: in "dog bites man", dog's query (1,0) scores the three keys (1,0), (0,1), (1,1) as 1, 0, 1; divided by √2 that is 0.707, 0, 0.707. Softmax gives weights 0.401, 0.198, 0.401, so dog's output is 0.401·(1,0) + 0.198·(0,1) + 0.401·(1,1) = (0.802, 0.599). In "man bites dog" dog sits in row 3, but it scores the same three keys in a different order, gets the same weights, and outputs the same (0.802, 0.599).

Level 3: in Python
import math
# q = k = v = the word vector, no positions
def Attn(X):
    out = []
    for q in X:
        scores = [sum(a * b for a, b in zip(q, k)) / math.sqrt(2) for k in X]
        exps = [math.exp(s) for s in scores]
        weights = [e / sum(exps) for e in exps]
        out.append(tuple(round(sum(w * v[c] for w, v in zip(weights, X)), 3) for c in range(2)))
    return out
# dog, bites, man
X = [(1, 0), (0, 1), (1, 1)]
# P swaps rows 1 and 3: man, bites, dog
PX = [X[2], X[1], X[0]]
Attn(X)  # → [(0.802, 0.599), (0.599, 0.802), (0.752, 0.752)]
# the same rows, shuffled the same way
Attn(PX)  # → [(0.752, 0.752), (0.599, 0.802), (0.802, 0.599)]

Shuffle the input and you get the same outputs, shuffled the same way. encode_sentence(..., scheme="none") runs a real attention layer and shows the "dog" row is the same vector whether "dog" comes first or last.

Figure 1 · Drawn from the lesson's code

no positions sinusoidal RoPE causal mask only 0.5 0.6 0.7 0.8 0.9 1.0 cosine similarity of pooled outputs 'dog bites man' vs 'man bites dog' 1.000 0.995 0.997 0.524

Without positions the two sentences are identical (similarity 1.0); sinusoidal codes and RoPE pull it just below 1, and a causal mask alone drops it to 0.52

Reading it: each bar compares the pooled attention output of "dog bites man" with that of "man bites dog", as a cosine similarity (1.0 means identical). With no positions the bar reaches exactly 1: the model literally cannot tell the sentences apart. Adding sinusoidal codes or RoPE pulls the bar below 1: the sentences now look different. The causal-mask bar is below 1 too, because a mask that lets token i see only i+1 tokens already leaks some order.

In code: word_vector gives each word a fixed random embedding, and order_similarity pools each sentence's outputs and compares them, which gives the bars in the figure.

Why it matters Almost all meaning in language is carried by order: negation scope, who did what to whom, code. So position must be injected. There are three families, below.

Chapter 1

Sinusoidal codes (the original 2017 transformer)

The clock with many hands, with two positions in your hands.

Everyday picture A clock with many hands. The second hand moves fast and tells nearby moments apart; the hour hand moves slowly and tells morning from evening. Read all the hands together and every moment has a unique fingerprint. Sinusoidal codes give every position such a set of "hands" and add them to the token's embedding.

Tiny worked example With 4 dimensions there are two hands. The fast one turns 1 radian per position; the slow one turns 1/100 of a radian:

position sin(pos·1) cos(pos·1) sin(pos/100) cos(pos/100)
0 0.000 1.000 0.000 1.000
1 0.841 0.540 0.010 1.000
2 0.909 −0.416 0.020 1.000

The fast pair already looks very different at positions 1 and 2; the slow pair barely moves, and it will only distinguish positions hundreds apart.

Figure 4 · Diagram

Reading it: two independent lookups meet at the plus sign. The top path says what the token is, the bottom path says where it is, and their sum is a single vector carrying both. Everything downstream sees only that sum, so position becomes part of the content attention compares.

The math and the code Two functions do the work. Picture a point walking around a circle of radius 1. The angle it has turned is measured in radians (a full turn is 2π ≈ 6.28 radians). Cosine of the angle is how far right the point is, and sine is how far up; both always lie between −1 and +1, and both repeat every full turn.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the token's position, counting from 0 1
how many numbers each position code has 4
which (sin, cos) pair, counting from 0 up to 0 and 1
, the two columns that pair fills: even column gets sine, odd gets cosine pair 1 fills columns 2 and 3
pair 's speed in radians per position (its frequency) ,
10,000 raised to a fraction; makes each pair slower than the last
, how far up / right a point is after turning by that many radians
the number in row , column of the code table

In words: "for position pos, each pair of columns stores where a hand turning at that pair's speed points after pos steps: sine in the first column, cosine in the second."

With the numbers: d = 4, pos = 1. Pair 0 turns 1 radian: sin 1 = 0.841, cos 1 = 0.540. Pair 1 turns 1/100 radian: sin 0.01 = 0.010, cos 0.01 = 1.000. So position 1's code is (0.841, 0.540, 0.010, 1.000), the second row of the table above.

Level 3: in Python
import math
d, pos = 4, 1
PE = []
# one (sin, cos) pair per i
for i in range(d // 2):
    # ω_i: pair i's speed
    omega_i = 1 / 10000 ** (2 * i / d)
    # columns 2i and 2i+1
    PE += [math.sin(pos * omega_i), math.cos(pos * omega_i)]
[f"{v:.3f}" for v in PE]  # → ['0.841', '0.540', '0.010', '1.000']

sinusoidal_encoding(n, d) builds the whole (n × d) table in four lines. A shift of k positions rotates every (sin, cos) pair by the same angle k·ω_i wherever you start, so the dot product of two position codes depends only on their distance.

Figure 3 · Drawn from the lesson's code

0 10 20 30 40 50 60 dimension (left: fast hands, right: slow hands) 0 25 50 75 100 125 150 175 position Sinusoidal position codes −1.00 −0.75 −0.50 −0.25 0.00 0.25 0.50 0.75 1.00 value

Left columns flip colour every few positions and right columns change slowly, so every row is a unique fingerprint and neighbours look alike

Reading it: each row is a position (0 at the top), each column one dimension, and colour is the value from −1 to +1. The left columns flip colour every few rows (fast hands); the right columns change slowly over hundreds of rows (slow hands). No two rows are identical, so every position has a unique fingerprint, and neighbouring rows look alike, so nearby positions get similar codes.

Why it matters Fixed codes need no training and exist for any position, and the distance property lets the model learn "look 3 tokens back" once and use it everywhere. But the code is mixed into the content vector itself, which later methods avoid.

Chapter 2

Learned absolute positions (GPT-2, BERT)

Fill the seats yourself, then read why the table can't grow.

Everyday picture Numbered theatre seats. Every seat has its own brass plate, and the theatre has a fixed number of seats.

Tiny worked example GPT-2 has a table of 1,024 position vectors, each 768 numbers long: 786,432 extra parameters, learned like any other weight. Position 5 looks up row 5. Position 1,024 has no row at all.

Figure 5 · Diagram

Reading it: identical to the sinusoidal picture except that the bottom path is a second trainable table instead of a formula. The dotted arrow is the catch: the table ends at max_len, so a longer input simply has nowhere to look.

The math and the code

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example (GPT-2)
the token's row of the token table (what the word means) 768 numbers
row of the position table, learned during training (where it is) 768 numbers; from 0 to 1,023
add the two lists number by number
the vector the first transformer block receives 768 numbers

In words: "the input vector is the word's meaning plus its seat number's learned vector."

With the numbers: with 2 dimensions, if = (0.5, −0.2) and = (0.1, 0.3), then = (0.6, 0.1). does not exist.

Level 3: in Python
E_dog = [0.5, -0.2]
# a position table with rows 0 to 3
P = [[0.0, 0.0], [0.0, 0.0], [0.0, 0.0], [0.1, 0.3]]
# x_3 = E_dog + P_3
[round(e + p, 2) for e, p in zip(E_dog, P[3])]  # → [0.6, 0.1]
# no row 1024: it cannot be encoded
len(P) > 1024  # → False

primer.ml.transformer.TinyGPT uses exactly this.

Why it matters It's simple and works well inside the trained length, but it cannot extrapolate. That limitation pushed the field to RoPE.

Chapter 3

Rotary position embeddings, RoPE (Llama, Mistral, Qwen and most modern LLMs)

Turn the query and the key yourself first; the rotation will then read as what it is.

Everyday picture Two people point at a clock face, one at 1 o'clock and one at 4 o'clock. The angle between their arms is 90°. Move both of them on to 8 and 11 o'clock and the angle is still 90°. RoPE turns each token's query and key like clock hands, by an amount proportional to its position. When a query meets a key, only the angle between them (the distance between the tokens) affects the score, not where on the clock they started.

Tiny worked example Use 2 dimensions, so there is one pair turning 1 radian per position. Let q = k = (1, 0). Put the query at position 3 and the key at position 7:

  • q turns by 3 rad: (cos 3, sin 3) = (−0.990, 0.141)
  • k turns by 7 rad: (cos 7, sin 7) = (0.754, 0.657)
  • score = −0.990·0.754 + 0.141·0.657 = −0.654 = cos(4)

Now move both 100 positions later, to 103 and 107. The score is still cos(107 − 103) = cos 4 = −0.654. Only the distance, 4, survived.

Figure 8 · Diagram

Reading it: this is ordinary attention with one extra step on the query and key branches. After projection, each (even, odd) pair of numbers in q and k is spun like a clock hand by an angle that grows with the token's position: fast for the first pairs, slow for the last. In the dot product only the difference of the angles survives, so the score knows how far apart the tokens are but not where they sit in the document. The value branch is untouched, so what gets blended is plain content, and nothing is ever added to the embedding.

The math and the code A rotation turns a point around the centre by some angle without changing its distance from the centre. Written as a 2 × 2 grid times a column of two numbers (a matrix-vector multiply: each output number is one row of the grid multiplied position-by-position with the input and added up):

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
pair of a query or key vector, before rotating (1, 0)
the same pair after rotating (−0.990, 0.141)
the token's position 3
which pair, from 0 to 0 (the only pair when d = 2)
pair 's turning speed in radians per position; means
the total angle turned 3 radians
the 2 × 2 grid the rotation: row 1 gives the new first number, row 2 the new second

In words: "new first number = cos(angle) × old first − sin(angle) × old second; new second number = sin(angle) × old first + cos(angle) × old second, where the angle is position times the pair's speed."

With the numbers: x = (1, 0) at m = 3, θ = 1: x′ = (cos 3 · 1 − sin 3 · 0, sin 3 · 1 + cos 3 · 0) = (−0.990, 0.141).

Level 3: in Python
import math
def rotate(x, m, theta_i=1.0):
    # the angle m θ_i
    a = m * theta_i
    # row 1 of the grid
    return (math.cos(a) * x[0] - math.sin(a) * x[1],
            # row 2 of the grid
            math.sin(a) * x[0] + math.cos(a) * x[1])
[f"{v:.3f}" for v in rotate((1, 0), m=3)]  # → ['-0.990', '0.141']
d, i = 2, 0
# θ_0 = 10000^(-2i/d)
10000 ** (-2 * i / d)  # → 1.0

Why only the distance survives: turning q by angle a and k by angle b and then taking their dot product (multiply matching numbers, add them up; large when the vectors point the same way) gives the same answer as turning k alone by b − a. So

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
, a query and a key vector, before rotating (1, 0) and (1, 0)
, the query's and the key's positions 3 and 7
"rotate every pair by its speed times " turn by 3 radians
the rotation for the distance only turn by 4 radians
the dot product of and (same as )

In words: "the score between a query at m and a key at n equals the score you'd get by leaving the query alone and turning the key by the gap n − m."

With the numbers: left side (−0.990, 0.141) · (0.754, 0.657) = −0.654; right side (1, 0) · (cos 4, sin 4) = cos 4 = −0.654. Same number, and the same again at positions 103 and 107.

Level 3: in Python
import math
# turn a pair by an angle (θ = 1, so the angle is the position)
def R(angle, x):
    return (math.cos(angle) * x[0] - math.sin(angle) * x[1],
            math.sin(angle) * x[0] + math.cos(angle) * x[1])
def dot(a, b): return sum(a_i * b_i for a_i, b_i in zip(a, b))
q = k = (1, 0)
# ⟨R_m q, R_n k⟩ with m = 3, n = 7
round(dot(R(3, q), R(7, k)), 3)  # → -0.654
# ⟨q, R_(n-m) k⟩: only the distance
round(dot(q, R(7 - 3, k)), 3)  # → -0.654
# 100 positions later, the same score
round(dot(R(103, q), R(107, k)), 3)  # → -0.654

apply_rope does this for a vector or a whole sequence with three element-wise lines, and no matrix multiply.

Figure 6 · Drawn from the lesson's code

0 10 20 30 40 50 60 distance between key and query (n − m) −12 −10 −8 −6 −4 −2 attention score q·k With RoPE the score depends only on distance query at position 0 query at position 100 query at position 1000

The score curves for starting positions 0, 100 and 1,000 lie exactly on top of each other: only the distance matters

Reading it: the x-axis is the distance n − m between a query and a key; each line uses the same q and k but places the pair at a different absolute starting position (0, 100, 1,000). The lines lie exactly on top of each other: the score depends only on distance. That is the relative-position property in one picture.

Figure 7 · Drawn from the lesson's code

0 25 50 75 100 125 150 175 200 position 0 1 2 3 4 5 6 rotation angle (mod 2π) RoPE: fast hands for local order, slow hands for long range pair 0 (θ=1) pair 4 (θ=0.32) pair 12 (θ=0.032) pair 31 (θ=0.00013)

Pair 0 wraps a full turn about every 6 tokens while the last pairs barely turn across the whole range

Reading it: each line is one pair of dimensions; the y-axis is how far that pair has turned (wrapped to one full turn, 2π) at each position on the x-axis. Pair 0 spins fastest and wraps every ~6 tokens, giving fine-grained local order. The last pairs barely turn across the whole range, giving coarse long-range position. It's the many-handed clock again, now applied inside attention.

In code: rope_frequencies returns each pair's turning speed θ_i, which apply_rope multiplies by the position to get every angle.

Why it matters Relative distance is what language actually uses ("the word two back"), rotation keeps vector lengths unchanged, and because position lives in angles, you can stretch it after training. That last point is the next section.

Chapter 4

Extending context after training

Everyday picture A ruler marked from 0 to 4,000. To measure something 32,000 long you can shrink the object by 8× so it fits on the ruler (interpolation), or re-draw only the long-distance marks while keeping the fine millimetre marks (NTK-aware scaling).

Tiny worked example A model trained on 4,096 tokens now reads 32,768. Interpolation multiplies every position by 4096/32768 = 1/8, so position 32,000 is treated as 4,000: an angle the model has seen. Neighbours are now 1/8 of a position apart, which a short fine-tune teaches it to resolve.

Figure 10 · Diagram

Reading it: both paths solve the same problem: angles beyond what the model saw in training. Interpolation slows every hand equally, which also blurs the fast hands that encode local word order. NTK-aware scaling slows only the slow hands and leaves the fastest one untouched, so local order stays sharp. Both usually end with a brief fine-tune.

The math and the code Position interpolation (interpolate_positions):

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the real position in the long input 32,000
the longest context seen in training 4,096
the context we now want 32,768
the position RoPE is actually given 4,000

In words: "shrink every position by the ratio of old length to new length."

With the numbers: 32,000 × 4,096 / 32,768 = 32,000 / 8 = 4,000.

Level 3: in Python
pos, L_train, L_new = 32_000, 4_096, 32_768
# pos' = pos · L_train / L_new
pos * L_train / L_new  # → 4000.0

NTK-aware scaling (ntk_scaled_base) changes the base instead:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the 10,000 inside 10,000
how many times longer the new context is 8
the head width 128
raised to a power just above 1
the new, larger base ≈ 82,685

In words: "multiply the base by slightly more than the stretch factor."

With the numbers: the fastest pair keeps radian per position. The slowest pair goes from $10000^{-126/128} = 1.15 \times 10^{-4}82685^{-126/128} = 1.44 \times 10^{-5}$: exactly 8× slower.

Level 3: in Python
base, s, d = 10_000, 8, 128
# s^(d/(d-2)): just above 8
round(s ** (d / (d - 2)), 2)  # → 8.27
# base' = base · s^(d/(d-2))
base_new = base * s ** (d / (d - 2))
round(base_new)  # → 82685
# θ_i for the last pair, 2i = 126
slowest = lambda b: b ** (-126 / 128)
f"{slowest(base):.2e}", f"{slowest(base_new):.2e}"  # → ('1.15e-04', '1.44e-05')
# exactly 8× slower
round(slowest(base) / slowest(base_new), 6)  # → 8.0

YaRN refines this per frequency band and is used by many long-context models. ALiBi skips position vectors entirely and subtracts a penalty proportional to distance from each attention score.

Figure 9 · Drawn from the lesson's code

0 5000 10000 15000 20000 25000 30000 position 0.0 0.5 1.0 1.5 2.0 2.5 3.0 3.5 angle of the slowest RoPE pair (rad) Keeping long contexts inside the trained range angles seen in training raw positions (extended) position interpolation NTK-aware base

Raw positions leave the trained angle band right after 4k and reach 8 times its top by 32k; interpolation and NTK scaling stay inside it

Reading it: the x-axis is position up to 32k and the y-axis is the angle of the slowest RoPE pair. The shaded band is the range of angles the model saw in training (positions up to 4k). The raw extended line leaves the band almost immediately after 4k: unfamiliar territory. The interpolated and NTK-scaled lines stay inside the band all the way to 32k.

Why it matters It's how models trained on a few thousand tokens are cheaply stretched to 100k or more, instead of being pretrained again from scratch.

The chart above follows the slowest pair. Here it is beside the fastest, with every pair in between.

Opening

A subtlety: the causal mask leaks order

In a decoder, token i sees exactly i+1 tokens, so even with no position encoding ("NoPE") the model can partly infer where it is. Bidirectional attention (BERT-style) has no such crutch and is fully order-blind.

In code: encode_sentence with its causal flag set runs the same layer under primer.ml.attention.causal_mask; its pooled similarity is the "causal mask only" bar in the first figure.

Test yourself

6 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Why does a transformer need positional information at all?Think it through, then reveal

Attention compares tokens by content only, so it's permutation-equivariant: shuffling the input just shuffles the output. Without positions, word order is invisible.

Question 2How does RoPE encode position, and what property makes it attractive?Think it through, then reveal

It rotates each pair of dimensions in q and k by an angle proportional to position. The q·k score then depends only on the relative distance, and rotation preserves vector length.

Question 3Compute a RoPE score by hand: 2 dimensions, θ = 1, q = k = (1, 0), query at 3, key at 7.Think it through, then reveal

cos(7 − 3) = cos 4 ≈ −0.654, and the same at positions 103 and 107.

Question 4Why do sinusoidal codes use many frequencies?Think it through, then reveal

Fast frequencies distinguish neighbours; slow ones distinguish distant positions. Together every position gets a unique, smoothly varying code, and a shift by k is the same rotation everywhere.

Question 5What goes wrong beyond the trained context length, and how do you fix it?Think it through, then reveal

The model meets position angles or indices it never saw, and attention degrades. Position interpolation, NTK-aware scaling or YaRN rescale positions or frequencies into the trained range, usually followed by a short fine-tune.

Question 6Learned absolute position embeddings: pro and con?Think it through, then reveal

They are simple and effective within the trained length, but they cannot represent any position beyond the table size.

Primary sources

The papers behind this lesson

Vaswani et al. (2017), Attention Is All You Need.

Introduced the transformer and, in section 3.5, the sinusoidal position codes built here.

Read on rumblr →The paper ↗
Su et al. (2021), RoFormer: Enhanced Transformer with Rotary Position Embedding.

Introduced RoPE: rotate queries and keys so attention scores depend only on relative distance.

Read the annotated companion →The paper ↗
Chen et al. (2023), Extending Context Window of Large Language Models via Positional Interpolation.

Showed that scaling positions down plus a short fine-tune extends RoPE models to much longer contexts.

The paper ↗
Peng et al. (2023), YaRN.

Refined RoPE context extension by treating fast and slow frequencies differently.

The paper ↗
Press et al. (2021), ALiBi.

Replaced position vectors with a distance penalty on attention scores that extrapolates to longer inputs.

The paper ↗

Researcher's shelf

Further reading

  • Vaswani et al., Attention Is All You Need (sinusoidal codes, section 3.5): https://arxiv.org/abs/1706.03762
  • Su et al., RoFormer: Enhanced Transformer with Rotary Position Embedding (2021): https://arxiv.org/abs/2104.09864
  • EleutherAI, Rotary Embeddings: A Relative Revolution: https://blog.eleuther.ai/rotary-embeddings/
  • Chen et al., Extending Context Window of Large Language Models via Positional Interpolation (2023): https://arxiv.org/abs/2306.15595
  • Peng et al., YaRN: Efficient Context Window Extension of Large Language Models (2023): https://arxiv.org/abs/2309.00071
  • Press et al., Train Short, Test Long: Attention with Linear Biases (ALiBi) (2021): https://arxiv.org/abs/2108.12409
  • Kazemnejad et al., The Impact of Positional Encoding on Length Generalization in Transformers (NoPE, 2023): https://arxiv.org/abs/2305.19466

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.