The lesson in one minute
What you'll be able to explain
- Tokenize the prompt, look up a vector per token, add position, run the transformer blocks, score every vocabulary entry from the last position, softmax, sample, append, repeat.
- Temperature divides the scores before softmax: low is predictable, high is varied.
- The naive loop re-reads everything each step; the KV cache makes each step cost one token. Training is the same forward pass plus a next-token loss.
Level 1
The practitioner's guide
In one sentence
A language model is a next-token scorer run in a loop: the prompt is cut into tokens, each token becomes a vector, the vectors pass through a stack of transformer blocks, the last position is scored against the whole vocabulary, one token is drawn, appended, and the loop runs again until a stop token or a length limit.
When you need it
You need this picture whenever you set a parameter
you can't explain (temperature, top_p, max_tokens), read a bill that
counts input and output tokens differently, or debug an answer that was cut
off, repeats itself, or comes out as gibberish. The tell: you are adjusting
a sampling knob by trial and error, or estimating cost in words when the
meter counts tokens. This lesson's toy tokenizer turns "Reset your password"
into 5 tokens, and " password" is one of them while "Reset" is two, so a
word count is only an estimate. You don't need this lesson to write a good
prompt, and you don't need it to pick a model by its benchmark scores; you
need it the first time a model's behaviour has to be predicted rather than
observed.
Your options
The loop has a few knobs a caller can turn, from the cheapest to the most certain:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Leave sampling at its defaults | Draws each token in proportion to the model's probabilities (temperature 1) | The variety the model was trained to produce | Run-to-run variation | The API's defaults |
| Lower the temperature, down to greedy | Divides the scores before softmax; at 0 it takes the single most likely token | The likely answer more often: at temperature 0.5 this lesson's top token rises from 0.63 to 0.84 | Blander text, and repetition at 0 (Holtzman et al., 2019); still not identical runs | One request parameter |
| Cut the tail (top-k, top-p) | Keeps only the k most likely tokens, or the smallest set whose probabilities reach p | No draws from the long tail of near-zero tokens | A knob some hosted APIs have withdrawn; open-model servers keep it | One request parameter |
Bound the length (max_tokens, stop sequences) |
Ends the loop at a token limit or at a string you name | A ceiling on cost and latency per call | Truncated answers if the limit is too low; check the stop reason | One request parameter |
| Reuse the prefix (KV cache, prompt caching) | Keeps the work done on tokens that haven't changed | Each new token costs one position instead of a re-read of everything: 219 positions against 23,900 for 200 tokens after a 20-token prompt | Server memory; caching rules and prices that differ by vendor | The serving stack |
| Constrain the output (schema, grammar) | Forbids, at every step, any token that cannot lead to a valid shape | A parseable answer by construction | A compiled schema and a small check per token | The model server (primer.ml.structured_output) |
How to choose
Start from who reads the answer and how long it is.
- Extraction, classification, tool calls: low temperature, a schema where
the server offers one, and a
max_tokenssized to the answer. - Writing, brainstorming, dialogue: the default temperature, and a length bound that stops runaway output rather than shaping it.
- Long documents in the prompt: the input is read in one parallel pass, so it costs money more than time; cache the unchanged prefix when you send it repeatedly.
- Agents and multi-turn systems: every turn re-sends the whole history, so the prompt grows with the conversation; bound each answer and watch the stop reason, because a truncated tool call is a broken one.
- Whatever you pick, measure tokens in and out on your own traffic. Output tokens are produced one loop iteration at a time, so the cheapest way to make a call faster is to ask for less.
What it costs
Two meters run. Input tokens go through the model in
one parallel pass, so a long prompt costs money and memory more than
time. Output tokens are produced one per trip round the loop, so latency is
proportional to length: this lesson's naive loop reads the prompt plus
everything generated so far at every step (14 positions to produce 4 tokens
from a 2-token prompt), and a KV cache turns that into one position per
token. Context is a hard ceiling: the position table has a fixed number of
rows (64 in this toy; max_position_embeddings in a Hugging Face config,
where LlamaConfig defaults to 2048, and 4k in the Llama 2 paper), and
prompt plus answer must fit inside it. Quality is measured on the same loop: the training loss is
the average of −ln p for the real next token, and perplexity is e raised
to it. This lesson's untrained model scores 5.97 against 5.99 for a blind
guess over 400 tokens (perplexity 393), and training over trillions of
tokens (2.0T for Llama 2) is what pushes that number down.
What breaks
- Gibberish. Random weights give a nearly flat guess over the vocabulary, and the demo's untrained model continues "The cat" with byte fragments. In practice the same symptom comes from a tokenizer that doesn't match the model, so ids fetch the wrong rows of the table. Check the pairing before the weights.
- Cut-off answers. The loop stopped at
max_tokens, not at a stop token. Claude's API reportsend_turnwhen the model finished on its own and another stop reason when your limit or stop sequence ended it; treat anything but a natural end as incomplete. - Temperature 0 that still varies. Greedy picks the argmax, but Claude's reference says results are not fully deterministic even at 0.0, and the arithmetic on real serving hardware is why.
- Repetition. Always taking the most likely token produces loops of the same phrase; Holtzman et al. (2019) showed it and proposed nucleus sampling as the fix.
- A bill that surprised you. Cost is counted in tokens, not words, and a conversation re-sends its whole history every turn.
- A context error. Prompt plus requested output exceeded the position
table. Shorten the prompt or the
max_tokens, or retrieve less.
In the wild
The pipeline is the decoder-only recipe of GPT-2 (Radford
et al., 2019: byte-level BPE, learned positions, tied embeddings) that
GPT-3 (Brown et al., 2020) scaled until instructions in the prompt were
enough. Hugging Face's generate() exposes the loop's knobs in a
GenerationConfig (do_sample, otherwise greedy; temperature 1.0,
top_k 50, top_p 1.0, max_new_tokens, repetition_penalty,
num_beams), and its LlamaForCausalLM returns logits of shape (batch,
sequence, vocabulary) with a logits_to_keep option because, as its docs
put it, only the last token's logits are needed for generation. Claude's
Messages API offers temperature (0.0 to 1.0, default 1.0; models
released after Claude Opus 4.6 accept only 1.0), max_tokens,
stop_sequences and a stop_reason on every response, and lets you set
max_tokens to 0 to warm the prompt cache without generating. Karpathy's
nanoGPT is this lesson at full size, in a few hundred lines. The papers are
linked at the end of the lesson.
Go deeper
Level 2 traces "Reset your password" through every box with its shapes, works temperature by hand on three scores, counts the loop's positions, and measures the loss of a model that has learned nothing. If you only needed to set the knobs, you are done.
Level 2
How it works, from scratch
This lesson wires the real pieces from the other lessons into one working
pipeline: the tokenizer from primer.ml.tokenization, the transformer from
primer.ml.transformer, and a sampling loop. Keep this one picture in your
head; every other lesson zooms into one box of it.
Figure 1 · Diagram
flowchart LR A[Prompt text] --> B[Tokenizer<br/>text to IDs] B --> C[Embedding lookup<br/>IDs to vectors] C --> D[Add position info] D --> E[Transformer blocks<br/>repeated N times] E --> F[Output layer<br/>score per vocab token] F --> G[Softmax + sampling] G --> H[Next token] H -->|append and repeat| E
Chapter 1
Text to ids: the coat check
Everyday picture A coat check. You hand over a coat (a piece of text) and get back a numbered ticket (a token id). The model only ever handles the tickets.
Tiny worked example With this module's toy tokenizer, "Reset your
password" becomes 5 tickets: Re set y our password →
[346, 377, 309, 328, 291]. The common word " password" got a single
ticket; the rarer pieces got several.
Figure 2 · Diagram
flowchart LR T["'Reset your password'"] --> TK["tokenizer<br/>(learned kit of 400 pieces)"] TK --> I["[346, 377, 309, 328, 291]"]
The code tok.encode(prompt); how the kit is learned is the whole of
primer.ml.tokenization.
In code: build_pipeline trains the toy tokenizer (primer.ml.tokenization.ByteBPE) and builds an untrained primer.ml.transformer.TinyGPT sized to its vocabulary.
Why it matters Prompt length, price and context limits are all counted in these tickets.
Chapter 2
Ids to vectors: a lookup, not a computation
Everyday picture A dictionary where ticket number 291 opens to page 291, and each page holds a list of numbers describing that token. A second dictionary, indexed by seat number, describes where the token sits.
Tiny worked example The table has 400 rows (one per token) of 32 numbers. Id 291 fetches row 291. The 5 ids fetch 5 rows: a 5 × 32 grid. Row i of the position table (i = 0 to 4) is added to row i of that grid.
Figure 3 · Diagram
flowchart LR
I["ids (5)"] --> E["token table<br/>400 × 32"]
E --> X["5 × 32: what each token is"]
P["positions 0..4"] --> PT["position table<br/>64 × 32"]
PT --> Y["5 × 32: where each token is"]
X --> ADD(("+"))
Y --> ADD
ADD --> OUT["5 × 32 input to the blocks"]
primer.ml.embeddings).The math and the code
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| a token's position in the prompt, from 0 | 4 (the last token) | |
| the token id at position | = 291 | |
| the token embedding table | 400 × 32 | |
| row of that table | row 291: 32 numbers | |
| row of the position table | row 4: 32 numbers | |
| what the first block receives for position | 32 numbers |
In words: "each token's input vector is its token's row plus its position's row."
With the numbers: , one 32-number list plus
another. The code is model.wte[ids] + model.wpe[:len(ids)]. In miniature,
with 3-number rows instead of 32: = (0.2, −0.1, 0.5) and =
(0.1, 0.3, −0.2) give = (0.3, 0.2, 0.3).
Level 3: in Python
# row 291 of the token table (3 numbers, not 32)
E_291 = [0.2, -0.1, 0.5]
# row 4 of the position table
P_4 = [0.1, 0.3, -0.2]
# E_(t_i) + P_i, number by number
x_4 = [round(e + p, 2) for e, p in zip(E_291, P_4)]
x_4 # → [0.3, 0.2, 0.3]
In code: trace does both lookups and the add, and keeps every stage's result so you can inspect the grid before and after positions are mixed in.
Why it matters This is the only place a token's identity enters the model; every later step works on these vectors.
Chapter 3
The transformer blocks: rounds of meeting and desk work
Everyday picture The team from primer.ml.transformer: each round is a
meeting where every token listens to the others (attention), then desk work
where each token thinks alone (feed-forward). This toy runs 2 rounds; large
models run dozens.
Tiny worked example The 5 × 32 grid goes into block 1 and comes out 5 × 32; the same through block 2. By the end, the vector at the last position ("password") has absorbed information from "Re", "set", "y" and "our".
Figure 4 · Diagram
flowchart LR X["5 × 32"] --> B1["block 1<br/>meeting + desk work"] --> B2["block 2"] --> LN["final norm"] --> H["5 × 32<br/>context-aware vectors"]
The math and the code
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the 5 × 32 input grid from section 2 | ||
| the -th transformer block | = 2 | |
| "and so on, for every block in between" | ||
| the final layer norm | ||
| the context-aware vectors | 5 × 32 |
In words: "run the input through every block in turn, then normalize."
With the numbers: here N = 2, so h = LN(Block₂(Block₁(x))). The code is
the for block in model.blocks loop in trace. In miniature, with one
3-number vector and two stand-in blocks that each add an edit: x = (1, 2, 3),
Block₁ adds (0, 1, 0) and Block₂ adds (0, 0, 2), giving (1, 3, 5). The final
norm subtracts the mean (3) and divides by the spread (1.63), so
h = (−1.22, 0, 1.22).
Level 3: in Python
import statistics
x = [1.0, 2.0, 3.0]
# a stand-in edit
def block_1(v): return [v_i + d for v_i, d in zip(v, [0, 1, 0])]
def block_2(v): return [v_i + d for v_i, d in zip(v, [0, 0, 2])]
def LN(v):
mu, sigma = statistics.fmean(v), statistics.pstdev(v)
# centre, then rescale
return [round((v_i - mu) / sigma, 2) for v_i in v]
h = x
# Block_1 first, then Block_2 ... up to Block_N
for block in [block_1, block_2]:
h = block(h)
h # → [1.0, 3.0, 5.0]
LN(h) # → [-1.22, 0.0, 1.22]
Why it matters This is where nearly all the compute and all the "understanding" happen.
Chapter 4
Scores, softmax and temperature
Everyday picture A scoreboard with one line for every token in the vocabulary. Softmax turns the scores into shares of a pie. Temperature sets how adventurous the pick is: low temperature almost always takes the biggest slice; high temperature gives the small slices a real chance.
Tiny worked example Three candidate tokens scored 2.0, 1.0 and 0.5.
| temperature | p(A) | p(B) | p(C) |
|---|---|---|---|
| 0 (greedy) | 1.000 | 0.000 | 0.000 |
| 0.5 | 0.844 | 0.114 | 0.042 |
| 1 | 0.629 | 0.231 | 0.140 |
| 2 | 0.481 | 0.292 | 0.227 |
Figure 7 · Diagram
flowchart LR H["last row of h<br/>32 numbers"] --> S["× token tableᵀ<br/>400 scores"] S --> T["÷ temperature"] T --> SM["softmax<br/>400 probabilities"] SM --> PICK["draw one token"]
The math and the code Softmax raises e (≈ 2.718) to the power of each score, so bigger scores get disproportionately bigger shares, then divides by the total so the shares add up to 1:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the score of candidate token | = 2.0 | |
| the temperature | 0.5 | |
| e ≈ 2.718 raised to that power | = 54.6 | |
| number of candidates (the vocabulary size) | 3 here, 400 in the model | |
| add up over every candidate | ||
| probability that token comes next | 0.844 |
In words: "the chance of token i is e to the power of its score over the temperature, divided by the same quantity summed over all tokens."
With the numbers: T = 0.5 doubles every score to (4, 2, 1): e⁴ = 54.6,
e² = 7.39, e¹ = 2.72, total 64.7, so p(A) = 54.6 / 64.7 = 0.844
(next_token_probs). Temperature 0 is the limit case: all probability on the
top score.
Level 3: in Python
import math
z = [2.0, 1.0, 0.5]
T = 0.5
# e^(z_i / T) for each candidate
exps = [math.exp(z_i / T) for z_i in z]
[round(e, 2) for e in exps] # → [54.6, 7.39, 2.72]
# Σ_j e^(z_j / T)
total = sum(exps)
round(total, 1) # → 64.7
# p_i: each share of the total
[round(e / total, 3) for e in exps] # → [0.844, 0.114, 0.042]
Figure 5 · Drawn from the lesson's code
At T = 0.25 token A takes 98% of the probability; at T = 2 the three tokens share it 48%, 29% and 23%
Figure 6 · Drawn from the lesson's code
The untrained model's 12 favourites are random byte fragments, each only 1.25 to 1.4 times the uniform 1-in-400 share
Why it matters Use low temperature for extraction and tool calls, where you want the most likely answer, and higher temperature for creative writing. Temperature 0 reduces randomness but doesn't guarantee identical outputs on real serving hardware.
Chapter 5
The loop: append and repeat
Everyday picture Writing a sentence one word at a time, and re-reading everything you've written before choosing each new word.
Tiny worked example A 2-token prompt, 4 new tokens. Step 1 reads 2 tokens, step 2 reads 3, step 3 reads 4, step 4 reads 5: 14 token positions to produce 4 tokens. With a 20-token prompt and 200 new tokens, the naive loop reads 23,900 positions; a KV cache reads 219.
Figure 9 · Diagram
flowchart LR S["ids so far"] --> M["full forward pass<br/>over ALL ids"] M --> P["probabilities for the next id"] P --> D["draw one id"] D --> A["append it"] A -->|"not done"| S A -->|"stop token or length limit"| OUT["decode ids to text"]
primer.ml.inference) stores it once.The math Total positions processed without a cache:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| prompt length in tokens | 2 | |
| tokens to generate | 4 | |
| the step counter, from 0 to | 0, 1, 2, 3 | |
| tokens re-read at step | 2, 3, 4, 5 | |
| total positions run through the model | 14 |
In words: "each step re-reads the prompt plus everything generated so far; add that up over all steps."
With the numbers: 4 × 2 + (4 × 3) / 2 = 8 + 6 = 14, the count
generate reports as positions_processed.
Level 3: in Python
p, n = 2, 4
# Σ over t = 0 .. n-1 of (p + t): 2 + 3 + 4 + 5
sum(p + t for t in range(n)) # → 14
# the closed form gives the same count
n * p + n * (n - 1) // 2 # → 14
Figure 8 · Drawn from the lesson's code
Without a cache work grows with the square of output length, 23,900 positions at 200 tokens against 219 with a KV cache
In code: Generation holds the generated ids, the decoded text and the positions count; naive_vs_cached_work counts the positions processed with and without a KV cache for the figure.
Why it matters Output length drives latency and cost; this is why every serving system caches keys and values.
Chapter 6
Training: the same forward pass, plus a loss
Everyday picture A guessing game with instant feedback. Cover the next word, guess it, uncover it, and note how surprised you were. Training nudges every weight to make the surprise smaller next time.
Tiny worked example If the model gives the right next token probability 0.5, the loss is −ln 0.5 = 0.693. Our untrained model averages 5.97 on a real sentence, close to ln 400 = 5.99, the score of a blind uniform guess.
Figure 10 · Diagram
flowchart LR T["training text"] --> F["forward pass<br/>(this whole lesson)"] F --> P["probabilities at every position"] P --> L["loss: −ln p(actual next token)<br/>averaged over positions"] L --> B["backpropagation<br/>gradient for every weight"] B --> U["optimizer nudges weights"] U -->|next batch| F
primer.ml.neural_net) works out how each
weight contributed, and the optimizer (primer.ml.optimizers) adjusts them.
One pass over n tokens gives n − 1 guesses at once, in parallel, which is a
big reason transformers train fast.The math and the code A logarithm answers "to what power must I raise e to get this number?" For probabilities between 0 and 1 it is negative, so we flip the sign; ln 1 = 0 (no surprise) and ln of a tiny number is very negative (huge surprise).
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| tokens in the training text | 12 ("The model reads tokens, not words.") | |
| the token at position | ||
| probability the model gave the actual next token, having seen everything before it; the bar reads "given" | ≈ 1/400 when untrained | |
| natural logarithm | ln(1/400) = −5.99 | |
| the average over all guesses | ||
| the loss training pushes down | 5.97 |
In words: "for every position, take the log of the probability the model gave the true next token, average them, and flip the sign."
With the numbers: a uniform guess over 400 tokens gives every true token
p = 1/400, so each term is −ln(1/400) = ln 400 = 5.99. Our untrained model
scores 5.97, so it is still essentially guessing. Perplexity, e raised to
the loss, is 393: "as unsure as choosing among 393 equally likely tokens"
(next_token_loss, cross_entropy).
Level 3: in Python
import math
# one guess that gave the true token p = 0.5
round(-math.log(0.5), 3) # → 0.693
n = 12
# a uniform guess gives every true token 1/400
p = [1 / 400] * (n - 1)
# -(1/(n-1)) Σ ln p
L = -sum(math.log(p_i) for p_i in p) / (n - 1)
round(L, 2) # → 5.99
Why it matters Pretraining is exactly this, over trillions of tokens: predicting the next token well forces the model to absorb grammar, facts and reasoning patterns. Random weights produce gibberish, and training is what turns the same machinery into a useful model.
Test yourself
6 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1What happens, step by step, when you send a prompt?Think it through, then reveal
Tokenizer turns text into ids; each id looks up an embedding; position information is added; the vectors pass through N transformer blocks; the last position's vector is scored against the whole vocabulary; softmax and sampling pick a token; it's appended and the loop repeats until a stop token.
Question 2Why does only the last position matter when generating?Think it through, then reveal
Its vector has attended to every earlier token and is the one trained to predict what comes next; earlier rows predict tokens we already have.
Question 3What does temperature do to the scores (2, 1, 0.5) at T = 0.5?Think it through, then reveal
Doubles them to (4, 2, 1) before softmax, sharpening the distribution: the top token rises from 0.63 to 0.84.
Question 4How many token positions does the naive loop process for a 2-token prompt and 4 new tokens?Think it through, then reveal
2 + 3 + 4 + 5 = 14. A KV cache avoids re-reading the unchanged prefix.
Question 5What loss does an untrained model get over a 400-token vocabulary, and why?Think it through, then reveal
About ln 400 ≈ 5.99 (perplexity about 400), because with random, tiny weights its predictions are close to uniform.
Question 6How is training different from inference?Think it through, then reveal
Same forward pass, plus a loss comparing predictions to the real next tokens, then backpropagation and a weight update. Inference only runs the forward pass.
Primary sources
The papers behind this lesson
The architecture inside the "transformer blocks" box.
Read on rumblr →The paper ↗The decoder-only, next-token-prediction recipe this pipeline follows, including learned positions, byte-level BPE and tied embeddings.
The paper ↗Showed that scaling this same loop up produces models that follow instructions from examples in the prompt.
Read the annotated companion →The paper ↗Introduced nucleus (top-p) sampling and explained why pure greedy decoding produces repetitive text.
The paper ↗Researcher's shelf
Further reading
- Andrej Karpathy, Let's build GPT (video): https://www.youtube.com/watch?v=kCc8FmEb1nY
- Karpathy's
nanoGPT: https://github.com/karpathy/nanoGPT - Jay Alammar, The Illustrated GPT-2: https://jalammar.github.io/illustrated-gpt2/
- 3Blue1Brown, But what is a GPT? (video): https://www.youtube.com/watch?v=wjZofJX0v4M
- Hugging Face, How to generate text (decoding strategies): https://huggingface.co/blog/how-to-generate
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.