At a glance
Key takeaways
- One idea: every modality becomes a sequence of vectors as wide as the language model's word vectors; after that it is ordinary attention.
- Images: a Vision Transformer cuts a picture into patches, projects each one, adds a position, and runs non-causal encoder blocks. Tokens = (H/P)·(W/P), so they grow with the square of the resolution.
- Connecting to a language model: a small projector maps image vectors into the embedding space, and they are spliced into the prompt. Train the projector first with everything frozen, then fine-tune on image instructions.
- Audio: waveform, then a log-mel spectrogram (a Fourier transform per 25 ms slice, pooled into ear-shaped bands), then an encoder: about 50 tokens a second.
- Generating pictures and sound: snap vectors to a learned codebook so they become ids a language model can predict like words; or hand off to a diffusion model.
- Video and cost: frames multiply image tokens by time; sample frames, merge tokens, and remember that text is by far the densest input.
Level 2
How it works, from scratch
Imagine a brilliant reader who can take in only one thing: a long row of
index cards, each card holding a short list of numbers. That is a language
model. Every word it reads arrives as one card, the word's vector (see
primer.ml.tokenization and primer.ml.transformer). It has never seen a
photograph or heard a voice.
To tell this reader about a photo, you cut the photo into small squares and write one card per square. To tell it about a voice recording, you write one card for every fiftieth of a second of sound. If the cards are written in the same "handwriting" as the word cards, the reader handles a photo the way it handles a sentence: it pays attention across all the cards at once.
That is the whole idea of a multimodal model, a model that takes in more than one modality (kind of input: text, images, audio, video):
Turn every modality into a sequence of vectors ("tokens") that a transformer can attend over.
Everything in this lesson is a way of making those cards: for images, for sound, for video, and, run backwards, for a model that produces pictures and speech.
Figure 1 · Diagram
flowchart LR IMG["Image"] --> VE["Vision encoder<br/>(patches to vectors)"] --> P1["Projector"] AUD["Audio"] --> SP["Spectrogram"] --> AE["Audio encoder"] --> P2["Projector"] TXT["Text"] --> TOK["Tokenizer"] --> EMB["Embedding table"] P1 --> SEQ["One sequence of vectors,<br/>all the same width"] P2 --> SEQ EMB --> SEQ SEQ --> LM["Decoder language model"] --> OUT["Next token:<br/>a word, or a codebook id<br/>for a picture or a sound"]
Chapter 1
A tiny worked example: a 4×4 picture becomes 4 tokens
Take a grayscale picture of 4×4 pixels, each pixel a brightness from 0 to 15:
0 1 | 2 3
4 5 | 6 7
------+------
8 9 | 10 11
12 13 | 14 15
- Cut it into 2×2 patches (the lines above): 4 patches.
- Flatten each patch into a list, reading row by row: the top-right patch becomes (2, 3, 6, 7).
- Project each list to the model's width, here 2 numbers, by multiplying by a matrix whose first column adds up the patch's top row and whose second column adds up its bottom row.
- Add a position: the patch's (row, column) in the grid, so the model knows where each patch came from.
| Patch | Pixels | Projected | + position | Token |
|---|---|---|---|---|
| top left | 0, 1, 4, 5 | (1, 9) | (0, 0) | (1, 9) |
| top right | 2, 3, 6, 7 | (5, 13) | (0, 1) | (5, 14) |
| bottom left | 8, 9, 12, 13 | (17, 25) | (1, 0) | (18, 25) |
| bottom right | 10, 11, 14, 15 | (21, 29) | (1, 1) | (22, 30) |
The picture is now four tokens of two numbers each: exactly the shape of input a transformer reads. A real model does the same with bigger numbers: patches of 14 or 16 pixels, and hundreds or thousands of numbers per token.
In code: worked_example_patches runs these four steps on this picture and returns the patches, the projections and the tokens.
Chapter 2
Images: a Vision Transformer from scratch
primer.ml.cnn_rnn introduced the idea of patches as tokens. Here we build
the whole front end, the Vision Transformer (ViT), and count what it
costs.
Cutting and counting
Everyday picture Lay a sheet of graph paper over a photo and cut along every fourth line. You get a grid of small tiles, and you can hand them over one by one, left to right and top to bottom, like the words of a sentence.
Tiny example A 224×224 image in 16-pixel patches is a grid of 224 / 16 = 14 patches down and 14 across: 14 × 14 = 196 tokens. A 336×336 image in 14-pixel patches is 24 × 24 = 576 tokens.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the image's height and width, in pixels | 224, 224 | |
| the side of one square patch, in pixels | 16 | |
| how many patches fit down the image (must be a whole number, so images are resized first) | 14 | |
| how many patches fit across | 14 | |
| multiply | ||
| the number of tokens the image becomes | 196 |
Each token starts life as numbers, where is the number of colour channels (3 for red, green, blue): 16 · 16 · 3 = 768.
In words: "the number of image tokens is the number of patches down times the number of patches across."
With the numbers: (224 / 16) · (224 / 16) = 14 · 14 = 196; (336 / 14) · (336 / 14) = 24 · 24 = 576; the worked 4×4 picture in 2-pixel patches gives 2 · 2 = 4.
In Python:
H, W, P = 224, 224, 16
# H/P patches down, times W/P across
(H // P) * (W // P) # → 196
H, W, P = 336, 336, 14
(H // P) * (W // P) # → 576
# the worked 4×4 picture in 2-pixel patches
(4 // 2) * (4 // 2) # → 4
Figure 2 · Drawn from the lesson's code
A 32 by 32 picture of a sun over a striped field, cut into 16 numbered 8-pixel patches, and the 16-row matrix those patches become
Figure 3 · Drawn from the lesson's code
Tokens per image rise with the square of the side length: 576 at 336 pixels, over 9,000 at 1,344 pixels with 14-pixel patches, and a quarter of that when 2 by 2 neighbours are merged
In code: image_to_patches cuts and flattens an image (the cutting is primer.ml.cnn_rnn.patchify), refusing sizes the patch does not divide, and count_image_tokens is the formula above, with an optional merge.
Why it matters in practice. Resolution is a cost dial. Small text in a
screenshot needs a high resolution to be legible, and every doubling of the
side costs four times the tokens (and attention cost grows faster still; see
primer.ml.attention). Systems resize images to a supported size, cut very
large ones into tiles, and merge neighbouring patches to keep the count
manageable.
From patch to token: projection and position
Everyday picture Every tile gets the same questionnaire: "how bright is your top half? your bottom half? is there an edge?" The answers become the tile's card. Then a sticker with the tile's grid address goes on the card, so shuffling the cards loses nothing.
Tiny example In the worked example, the questionnaire had two questions (top-row total, bottom-row total), and the sticker was the patch's (row, column).
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape | In the example (top-right patch) |
|---|---|---|---|
| which patch, counting in reading order from 0 | 1 | ||
| patch flattened into a row of pixels | (2, 3, 6, 7) | ||
| the learned patch-embedding matrix: one column per output number (a bias vector is usually added too; it is left out here) | the 4 × 2 matrix below | ||
| a matrix multiply: each output number is the dot product of the patch with one column of | (5, 13) | ||
| the learned position vector for slot | (0, 1) | ||
| the finished token for patch | (5, 14) | ||
| the model width: how many numbers per token | 2 |
in the example is the rows (1, 0), (1, 0), (0, 1), (0, 1): the first two pixels (the patch's top row) feed output 1, the last two (its bottom row) feed output 2.
In words: "each token is its patch multiplied by a learned matrix, plus a learned vector that marks where the patch sits."
With the numbers: (2, 3, 6, 7) · = (2 + 3, 6 + 7) = (5, 13); adding the position (0, 1) gives (5, 14).
In Python:
x_i = [2, 3, 6, 7]
W_E = [[1, 0], [1, 0], [0, 1], [0, 1]]
p_i = [0, 1]
# x_i W_E: dot the patch with each column of W_E
xW = [sum(x * W_E[r][c] for r, x in enumerate(x_i)) for c in range(2)]
xW # → [5, 13]
# + p_i
[a + b for a, b in zip(xW, p_i)] # → [5, 14]
Why the position vector? Attention compares every token with every other
and ignores their order (see primer.ml.positional). Without , a sky
patch at the top and the same sky colour in a puddle at the bottom would be
the same token, and "the sun is above the field" could not be expressed.
Figure 4 · Diagram
flowchart LR IMG["Image<br/>H × W × C"] --> CUT["Cut into N patches<br/>N × P·P·C"] CUT --> PROJ["Multiply by W_E<br/>N × d"] PROJ --> POS["Add position vectors<br/>N × d"] POS --> ENC["Encoder blocks<br/>every patch attends to every patch<br/>(no causal mask)"] ENC --> OUT["N image vectors<br/>N × d"]
primer.ml.transformer, run with the causal mask switched off. A sentence
is read left to right, so a text decoder hides the future; a picture has no
future, so every patch may look at every other, above, below and to either
side. The output has one vector per patch, now informed by the whole image.In code: VisionEncoder holds W_E, the position table and a stack of primer.ml.transformer.TransformerBlock with causal=False; VisionEncoder.embed is the formula above, and calling the encoder runs the blocks.
Why it matters in practice. The vision encoder inside most
vision-language models is a ViT that was first trained as the image half of
CLIP (see primer.ml.embeddings.contrastive). CLIP training pulled its
image vectors towards the vectors of matching captions, so its outputs
already carry meaning a language model can use. That is why builders start
from it instead of training vision from scratch.
Chapter 3
Connecting vision to a language model
The projector: a plug adapter
Everyday picture Your laptop charger has the right voltage but the wrong plug for the wall socket abroad. A small adapter changes the shape, not the power. The vision encoder speaks in vectors of its own width and style; the language model expects vectors shaped like its word embeddings. A projector is the adapter between them.
Tiny example A vision vector with 2 numbers, (1, 2), must become a language-model vector with 3 numbers.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape | In the example |
|---|---|---|---|
| one image token from the vision encoder | (1, 2) | ||
| the projector's learned matrix | rows (1, 0, 1) and (0, 1, 1) | ||
| the same token, now shaped like a word embedding | (1, 2, 3) | ||
| the widths of the two models | 2 and 3 |
In words: "multiply each image vector by one learned matrix to turn it into a vector the language model can read."
With the numbers: (1, 2) · = (1·1 + 2·0, 1·0 + 2·1, 1·1 + 2·1) = (1, 2, 3).
Level 3: in Python
v = [1, 2]
W_P = [[1, 0, 1], [0, 1, 1]]
# v W_P: dot v with each column of W_P
[sum(v[r] * W_P[r][c] for r in range(2)) for c in range(3)] # → [1, 2, 3]
Many models use two such layers with a GELU between them (a smooth
"keep the positives" function; see primer.ml.transformer), which is a
small feed-forward network. Either way, the projector is tiny next to the
two models it joins: a few million numbers between models with billions.
In code: Projector is a single linear layer by default and a two-layer MLP when given a hidden width; worked_example_projection computes the (1, 2, 3) above.
Splicing image tokens into the prompt
Everyday picture Writing a letter and taping a strip of photos into the middle of a sentence. The reader reads the words, then the photos, then the rest of the words, in order.
Tiny example The prompt "what is this <image>" has four text tokens,
one of them a placeholder. The placeholder is replaced by the image's 4
projected vectors, so the language model reads 3 + 4 = 7 rows.
| Step | Shape (toy model) | Meaning |
|---|---|---|
| image | (8, 8) | an 8×8 grayscale picture |
| patches | (4, 16) | four 4×4 patches, flattened |
| vision encoder output | (4, 16) | four image vectors of width 16 |
| projector output | (4, 24) | four image tokens, as wide as the language model |
| text token vectors | (3, 24) | "what", "is", "this" from the embedding table |
| spliced sequence | (7, 24) | text, then image, in prompt order |
| logits | (7, 12) | a score for each of the 12 vocabulary words, at every position |
Figure 5 · Diagram
flowchart LR IMG["Image"] --> VE["Vision encoder<br/>4 × 16"] --> PR["Projector<br/>4 × 24"] TXT["Prompt: what is this,<br/>then the image placeholder"] --> TE["Look up text tokens<br/>3 × 24"] PR --> SPL["Splice at the placeholder<br/>7 × 24"] TE --> SPL SPL --> POS["+ positions"] --> DEC["Decoder blocks<br/>(causal)"] --> LOG["Logits<br/>7 × 12"]
Because the decoder is causal, a token can use the image only if it comes after the image. Text before the picture is scored exactly as if there were no picture. That is why prompts usually put the image first and the question after it.
In code: VisionLanguageModel.splice builds the spliced sequence and a flag for each row saying whether it is an image token; calling VisionLanguageModel runs primer.ml.transformer.TinyGPT's blocks on it; build_toy_vlm wires up the toy model in the table.
Why it matters in practice. This "encode, project, splice" design (as in LLaVA) is the simplest and most common. Two other families exist. One adds new cross-attention layers inside the language model that look at the image vectors (as in Flamingo), leaving the text sequence short. Another first compresses any image into a fixed small number of tokens, such as 32 or 64, with a little attention module of learned queries (as in BLIP-2's Q-Former). Both trade some detail for fewer tokens.
How such a model is trained
Everyday picture Two experts who don't share a language: a photographer and a writer. You hire an interpreter. First, the interpreter learns vocabulary while both experts carry on exactly as they are: the photographer points at pictures, the writer names them. Then all three practise real conversations together, and the writer is allowed to adapt a little too.
That is the usual two-stage recipe:
- Alignment. Freeze the vision encoder and the language model. Train only the projector on image-caption pairs, so the projected image tokens make the language model produce the caption.
- Visual instruction tuning. Train the projector and the language model
(fully, or with LoRA; see
primer.ml.training_stages) on images paired with instructions and answers: "what is unusual about this picture?" followed by a good answer.
In both stages the loss is ordinary next-token cross-entropy (the
negative log of the probability given to the right token; see
primer.ml.losses), counted only on the answer tokens. The model is not
graded on predicting the image tokens or the question it was given.
Tiny example The answer is "horizontal stripes", two tokens. The model gives the right first token probability 0.5 and the right second token 0.25.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the positions of the answer tokens (the loss mask keeps only these) | 2 positions | |
| how many answer tokens there are | 2 | |
| "for each position in the answer" | ||
| the correct token at position | "horizontal", then "stripes" | |
| every token before position | ||
| the probability the model gives the correct token, given () the image and what came before | 0.5, 0.25 | |
the natural logarithm; is 0 when and grows as shrinks (see primer.notation) |
||
| the loss: the average surprise over the answer | 1.040 |
In words: "the loss is the average, over the answer tokens only, of how surprised the model was by each correct token."
With the numbers: −(log 0.5 + log 0.25) / 2 = (0.693 + 1.386) / 2 = 1.040.
Level 3: in Python
import math
# probabilities of the right answer tokens
p = [0.5, 0.25]
# −(1/|A|) Σ log p
round(-sum(math.log(q) for q in p) / len(p), 3) # → 1.04
Figure 6 · Diagram
flowchart TB
subgraph S1["Stage 1: alignment (image-caption pairs)"]
direction LR
V1["Vision encoder<br/>frozen"] --> P1["Projector<br/>TRAINED"] --> L1["Language model<br/>frozen"]
end
subgraph S2["Stage 2: visual instruction tuning (image, question, answer)"]
direction LR
V2["Vision encoder<br/>frozen"] --> P2["Projector<br/>trained"] --> L2["Language model<br/>trained or LoRA"]
end
S1 --> S2
Our toy does stage 1. Its vision encoder and language model are random and frozen; only the projector learns, from 80 small pictures of four patterns (horizontal, vertical, diagonal, checkered), to make the language model's output layer pick the right pattern word.
Figure 7 · Drawn from the lesson's code
Stage-1 alignment: the caption loss falls from 2.41 to 0.33, and on 80 new images the right word ranks first 99% of the time, up from 24%
In code: align_projector trains only the projector's weights with cross-entropy on the caption word and returns the loss and accuracy per step; caption_accuracy measures how often the right word wins. The toy scores a pooled image vector directly against the language model's token table (the last step of the real path) rather than back-propagating through every layer.
Why it matters in practice. Stage 1 is cheap enough to run in hours, because almost every weight is frozen. Stage 2 decides the model's behaviour: the quality and variety of the image-instruction data matter more than its size. A common failure of a weakly aligned model is describing objects that are not in the picture, a visual form of hallucination.
Chapter 4
Audio: from a waveform to tokens
Sound is a list of numbers
Everyday picture A microphone is a tiny eardrum. It measures air pressure many thousands of times a second, and the recording is just that list of measurements: the waveform. Drawn on paper, it looks like a seismograph trace.
Tiny example Speech is usually recorded at a sample rate of 16,000 measurements per second, so one second is 16,000 numbers and a 30-second clip is 480,000. Used directly as tokens that would be ruinous, and the thing that matters, which pitches are sounding, is hidden in the wiggles (top panel of the spectrogram figure below).
How much of each pitch: the Fourier transform
Everyday picture A prism splits white light into its colours. The Fourier transform splits a sound into its pitches. It works by holding a pure wave of each pitch up against the recording and asking "how well do you line up?". A pitch that is present lines up again and again and scores high; a pitch that is absent lines up as often as it clashes and scores zero.
Tiny example Eight samples of a wave that goes up and down twice: (1, 0, −1, 0, 1, 0, −1, 0). Line it up against a wave that also cycles twice, cos: (1, 0, −1, 0, 1, 0, −1, 0). Multiply matching positions and add: 1 + 0 + 1 + 0 + 1 + 0 + 1 + 0 = 4. A wave that cycles once agrees for half its length and disagrees for the other half: it scores 0.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the -th sample of the recording | (1, 0, −1, 0, 1, 0, −1, 0) | |
| how many samples | 8 | |
| a counter over the samples, from 0 to | 0, 1, …, 7 | |
| which pitch we are testing: a wave that completes cycles in the samples | 2 | |
| the cosine and sine waves (sine is the same wave shifted a quarter cycle, so a pitch that starts at a different moment is still caught) | ||
| one full cycle, in the units cos and sin use | ||
| add up over every sample | ||
| the length of the pair (cosine score, sine score) | ||
| how much of pitch the recording holds | 4 |
Bin is the frequency in hertz (cycles per second), where is the sample rate. If the 8 samples were taken over one second, and bin 2 is 2 Hz. Textbooks write the same thing with complex numbers, ; the cosine part and the sine part are exactly the two sums above.
In words: "for each pitch, correlate the recording with a cosine and a sine of that pitch, and take the length of the two scores."
With the numbers: for the cosine sum is 4 and the sine sum is 0, so . For both sums are 0. Bin 6 also scores 4: with only 8 samples, a wave cycling 6 times is indistinguishable from one cycling 8 − 6 = 2 times, so the top half of the bins mirrors the bottom half and only bins 0 to are kept.
In Python:
import math
x = [1, 0, -1, 0, 1, 0, -1, 0]
N = len(x)
def magnitude(k):
c = sum(x_n * math.cos(2 * math.pi * k * n / N) for n, x_n in enumerate(x))
s = sum(x_n * math.sin(2 * math.pi * k * n / N) for n, x_n in enumerate(x))
return math.sqrt(c ** 2 + s ** 2)
[round(magnitude(k), 6) for k in range(N)] # → [0.0, 0.0, 4.0, 0.0, 0.0, 0.0, 4.0, 0.0]
# f_k = k · f_s / N, with 8 samples a second
2 * 8 / N # → 2.0
In code: dft_magnitudes is this formula for every k at once, written from the definition (the tests check it against NumPy's fast Fourier transform, which computes the same numbers far faster).
The spectrogram: one Fourier transform per slice
Everyday picture A piano roll, or sheet music: time runs left to right, pitch runs bottom to top, and a mark means "this note is sounding now". To make one from a recording, listen through a short window, say 25 milliseconds, ask "which pitches are in here?", slide the window along a little, and ask again. That is the short-time Fourier transform (STFT), and its picture is a spectrogram.
Each slice is faded in and out with a Hann window (a smooth hump from 0 up to 1 and back) before measuring, because chopping a wave off abruptly creates a click, and a click contains every pitch at once.
Tiny example Speech systems commonly use 25 ms windows (400 samples at 16 kHz) that start every 10 ms (160 samples). How many windows fit in one second?
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the recording's length, in samples | 16,000 | |
| the window length, in samples | 400 (25 ms) | |
| the hop: how far the window moves each step | 160 (10 ms) | |
| floor: round down to a whole number (a part-window at the end is dropped) | ||
| the number of frames (columns of the spectrogram) | 98 |
In words: "one window fits at the start; after that, count how many whole hops still leave room for a full window."
With the numbers: 1 + ⌊(16,000 − 400) / 160⌋ = 1 + ⌊97.5⌋ = 98 frames per second, which is why speech models talk about "100 frames a second".
Level 3: in Python
L, N, H = 16000, 400, 160
# 1 + ⌊(L − N) / H⌋
1 + (L - N) // H # → 98
Figure 8 · Diagram
flowchart LR W["Waveform<br/>L samples"] --> F["Slice into frames<br/>T × N"] F --> HW["Fade each frame<br/>(Hann window)"] HW --> DFT["Fourier transform<br/>per frame<br/>T × (N/2 + 1)"] DFT --> MEL["Pool into mel bands<br/>T × n_mels"] MEL --> LOG["Take the log<br/>log-mel spectrogram"]
Figure 9 · Drawn from the lesson's code
A 20-millisecond waveform of wiggles, then a spectrogram of the same second of sound showing a flat line at 1000 Hz and a line rising from 200 to 3000 Hz, then the log-mel version where the rising line curves
In code: stft slices, windows (with hann) and measures every frame; frame_count is the formula above; log_mel_spectrogram adds the mel pooling and the log.
The mel scale: spacing pitch the way ears do
Everyday picture On a piano, every octave takes up the same width of keyboard, yet each octave doubles the frequency: the A keys are at 110, 220, 440 and 880 Hz. Ears work the same way above about 1,000 Hz: a jump from 1,000 to 2,000 Hz sounds about as big as a jump from 2,000 to 4,000. The mel scale relabels frequencies so that equal steps in mel sound like equal steps in pitch.
Tiny example By construction 1,000 Hz is 1,000 mel. The 7,000 Hz from 1,000 to 8,000 Hz shrinks to just 1,840 mel.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| a frequency in hertz | 700 | |
| frequency measured in units of 700 Hz; below about 700 Hz the scale is nearly straight, above it the log takes over | 1 | |
the base-10 logarithm: "10 to what power gives this?"; it turns ratios into equal steps (see primer.notation) |
||
| 2595 | a constant chosen so that 1,000 Hz comes out as 1,000 mel | |
| the same frequency in mel | 781.2 |
In words: "mel is a log of the frequency, gently straightened below 700 Hz and scaled so that 1,000 Hz is 1,000 mel."
With the numbers: 2595 · log₁₀(1 + 700/700) = 2595 · 0.301 = 781.2; 1,000 Hz gives 1,000.0; 4,000 Hz gives 2,146.1; 8,000 Hz gives 2,840.0.
Level 3: in Python
import math
def mel(f):
return 2595 * math.log10(1 + f / 700)
round(mel(700), 1) # → 781.2
round(mel(1000), 1) # → 1000.0
round(mel(4000), 1) # → 2146.1
A mel filterbank turns the hundreds of frequency bins into a few dozen bands: triangles evenly spaced in mel, so narrow at low frequencies and wide at high ones. Each band adds up the energy under its triangle. Loudness is also heard by ratio, so the last step takes the log.
Figure 10 · Drawn from the lesson's code
Left: the mel curve rises steeply to 1000 mel at 1000 Hz and then flattens, reaching only 2840 at 8000 Hz. Right: ten triangular filters, narrow below 1000 Hz and ever wider above
In code: hz_to_mel is the formula, and mel_filterbank builds the triangles, evenly spaced in mel.
Speech recognition: an encoder over frames, a decoder for text
Everyday picture A court stenographer listens to a whole sentence, then types it out word by word, glancing back at what they heard as they go.
Tiny example Take a 30-second clip. With 10 ms hops it is 3,000 frames of 80 mel bands each. A small convolution that moves two frames at a time halves that to 1,500 encoder positions: 50 audio tokens per second. A transformer encoder attends over all 1,500 at once, and a decoder writes the transcript as text tokens, attending both to its own words so far and to the encoder's output. This is the design of Whisper.
Figure 11 · Diagram
flowchart LR A["30 s of audio"] --> LM["Log-mel spectrogram<br/>3000 frames × 80 bands"] LM --> CV["Convolution, stride 2<br/>1500 positions"] CV --> ENC["Encoder blocks<br/>(no causal mask)"] ENC --> X["Cross-attention:<br/>decoder reads the audio"] DEC["Decoder blocks<br/>(causal)"] --> X X --> TXT["Text tokens,<br/>one at a time"] TXT -.-> DEC
primer.ml.transformer for encoders and decoders). The dotted arrow is the
generation loop: each new word is fed back in.In code: audio_tokens counts encoder positions from a clip's length, the hop and the downsampling.
Why it matters in practice. The same audio encoder can feed a general language model instead of a dedicated decoder: add a projector, splice the audio tokens into the prompt, and the model can answer questions about a recording, exactly as with images. Fifty tokens a second is manageable for a short voice message and expensive for an hour-long meeting.
Chapter 5
Generating images and speech: discrete tokens from a codebook
So far the model reads pictures and sound. To write them, a language model needs them as something it can predict one at a time from a fixed vocabulary, the way it predicts words.
Everyday picture A paint-by-numbers kit comes with a palette of numbered pots. Any small area of a picture is described by the number of the pot closest to its colour. The whole picture becomes a grid of pot numbers, and anyone with the same palette can repaint it. The palette is a codebook; snapping to the nearest entry is vector quantization.
Tiny example A codebook of three 2-number entries: , , . The vector has squared distances 0.85, 0.05 and 1.45 to them, so it becomes id 1. Decoding id 1 gives back (1, 0): close, not exact.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| one vector to encode: an image patch's vector or a slice of sound | (0.9, 0.2) | |
| codebook entry number | ||
| the codebook size: how many entries | 3 | |
| " ranges over the entry numbers" | 0, 1, 2 | |
| squared distance: subtract, square each difference, add | ||
| "the that gives the smallest value" (the position of the minimum, not the minimum itself) | 1 | |
| the id that replaces | 1 |
In words: "replace each vector by the number of its nearest codebook entry."
With the numbers: distances squared to are 0.81 + 0.04 = 0.85, 0.01 + 0.04 = 0.05 and 0.81 + 0.64 = 1.45; the smallest is 0.05, at .
Level 3: in Python
z = [0.9, 0.2]
codebook = [[0, 0], [1, 0], [0, 1]]
# ‖z − c_k‖² for each entry
d2 = [sum((a - b) ** 2 for a, b in zip(z, c)) for c in codebook]
[round(d, 2) for d in d2] # → [0.85, 0.05, 1.45]
# arg min: the position of the smallest
d2.index(min(d2)) # → 1
Figure 12 · Diagram
flowchart LR IN["Picture or sound"] --> E["Encoder<br/>vectors"] E --> Q["Snap to nearest<br/>codebook entry"] Q --> IDS["Ids, like word tokens<br/>e.g. 17, 803, 4, ..."] IDS --> LM["Language model<br/>learns to predict ids<br/>after text"] LM --> GEN["Generated ids"] GEN --> LOOK["Look up codebook<br/>vectors"] LOOK --> D["Decoder"] --> OUTP["Pixels or waveform"]
primer.ml.generative.autoencoders).How is the codebook chosen? It is learned so that the entries sit where the
data actually is, which amounts to k-means clustering (see
primer.ml.embeddings.clustering): snap every vector to its nearest entry,
move each entry to the average of the vectors that chose it, and repeat.
For audio, one codebook is rarely precise enough, so neural audio codecs
use residual vector quantization: a second codebook encodes what the
first one missed, a third what the second missed, and so on. Each moment
of sound becomes a small stack of ids.
Figure 13 · Drawn from the lesson's code
On new patches, a learned codebook's error falls from 0.32 to about 0.09 as it grows to 64 entries while random codebooks stay far above it; stacking residual codebooks of 8 entries keeps pushing the error down
In code: quantize is the arg min above and dequantize the lookup; kmeans_codebook learns a codebook; learn_residual_codebooks and residual_quantize stack codebooks on the leftovers.
Why it matters in practice. Discrete tokens let one model read and
write every modality with one mechanism, next-token prediction. The costs
are real: a 256×256 image compressed eight-fold per side is 32 × 32 =
1,024 ids (the scheme of the original DALL·E), and a neural audio codec
running at 75 frames a second with 8 codebooks writes 600 ids a second. The
other road to generating images is diffusion (see
primer.ml.generative.diffusion): the language model supplies text or
vectors that condition a diffusion model, which paints the pixels in many
small denoising steps. Discrete tokens are simpler to bolt onto a language
model; diffusion usually renders finer detail.
Chapter 6
Video: pictures over time
Everyday picture A flipbook. Each page is a picture, and neighbouring pages are almost identical. To follow the story you don't need to look at every page; a glance at every tenth page shows what happens.
Tiny example A 10-second clip at 30 frames a second, with each frame cut into 256 tokens (224×224 in 14-pixel patches): 300 frames × 256 = 76,800 tokens, more than half of a 128,000-token context, for ten seconds. Keeping one frame a second gives 10 × 256 = 2,560.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the clip's duration, in seconds | 10 | |
| frames kept per second (the sampling rate, often far below the video's own frame rate) | 30, then 1 | |
| the tubelet depth: how many consecutive frames share one token, when a patch is cut through time as well as space | 1 | |
| how many frame slots become tokens | 300, then 10 | |
| tokens per frame, from the image formula | 16 · 16 = 256 | |
| the clip's total tokens | 76,800, then 2,560 |
In words: "count the frames you keep, divide by how many frames each token spans, and multiply by the tokens per frame."
With the numbers: (10 · 30 / 1) · 16 · 16 = 300 · 256 = 76,800; at one frame a second, (10 · 1 / 1) · 256 = 2,560; sampling 2 frames a second in tubelets 2 frames deep gives (10 · 2 / 2) · 256 = 2,560 again, while covering twice as many moments.
Level 3: in Python
H = W = 224
P = 14
tokens_per_frame = (H // P) * (W // P)
tokens_per_frame # → 256
D, r, t = 10, 30, 1
D * r // t * tokens_per_frame # → 76800
# one frame a second
D, r, t = 10, 1, 1
D * r // t * tokens_per_frame # → 2560
Figure 14 · Diagram
flowchart LR V["Video<br/>30 frames a second"] --> S["Sample frames<br/>e.g. 1 a second"] S --> ENC["Vision encoder per frame<br/>(or tubelets across frames)"] ENC --> M["Merge neighbouring<br/>patch tokens"] M --> PR["Projector"] PR --> SEQ["Frame tokens,<br/>with timestamps as text"] SEQ --> LM["Language model"]
In code: video_tokens is the formula, and sample_frames picks the evenly spaced frames to keep.
Why it matters in practice. Sampling is a bet that nothing important happens between the frames you keep: one frame a second is fine for a lecture and useless for a golf swing. Real systems adapt: more frames for short or fast clips, fewer and smaller ones for long recordings, and a transcript of the soundtrack alongside.
Chapter 7
What it costs: tokens per second and the context budget
Everyday picture Packing a suitcase. Text is neatly folded shirts; images are shoes; video is an inflatable boat. The suitcase (the context window) is the same size whatever you pack.
Tiny example A 128,000-token window holds 128,000 / 256 = 500 seconds, about 8 minutes, of video sampled at one frame a second, before a single word of the question. The same window holds about 11 hours of speech written out as text.
| Input | Tokens per second | Minutes in 128k tokens |
|---|---|---|
| speech written as text (about 150 words a minute, about 1.3 tokens a word) | 3.2 | 656 |
| speech as audio-encoder tokens | 50 | 43 |
| video, 1 frame a second, 256 tokens a frame | 256 | 8.3 |
| audio as codec tokens (75 frames × 8 codebooks) | 600 | 3.6 |
| video, 30 frames a second, 256 tokens a frame | 7,680 | 0.3 |
Figure 15 · Drawn from the lesson's code
On log axes, five straight lines of tokens against recording length: speech-as-text crosses a 128k window only after hours, 30-frames-a-second video within seconds
In code: seconds_that_fit divides a window by a rate, and TOKENS_PER_SECOND holds the rates in the table.
Why it matters in practice. Every image and second of audio is billed
and computed like text: it takes room in the context window, lengthens the
prefill before the first word of the answer (see primer.ml.inference),
and competes with instructions and retrieved documents for attention (see
primer.agents.context). The practical levers follow from the formulas:
send the smallest resolution that still shows the detail you need, crop to
the region that matters, sample fewer frames, and transcribe audio to text
when the words matter and the tone does not.
Test yourself
8 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1How can a chatbot "see" a photo, explained without jargon?Think it through, then reveal
The photo is cut into a grid of small squares, and each square is described by a list of numbers in the same format the model uses for words. The model then reads the squares and the words of your question together, paying attention to whichever squares help answer it. It never sees the picture as a picture: it reads it as a few hundred extra "words".
Question 2How many tokens is a 448×448 image with 14-pixel patches? And if each 2×2 block of neighbours is merged?Think it through, then reveal
448 / 14 = 32 patches a side, so 32 × 32 = 1,024 tokens. Merging 2×2 blocks divides by 4: 256 tokens. Doubling the side from 224 would have quadrupled the count.
Question 3Why does a Vision Transformer need position vectors, and why is its attention not causal?Think it through, then reveal
Attention ignores order, so without positions the model sees an unordered bag of patches and cannot tell the sky from the ground. It is not causal because a picture has no "future": every patch may use every other, and hiding half the image would only throw information away.
Question 4What does the projector do, and why train it first with both big models frozen?Think it through, then reveal
It maps each image vector into the language model's embedding space, so image tokens look like something the language model can use. Training it alone first is cheap (it is tiny) and safe (the frozen models keep all they know). Only once the two sides understand each other is the language model tuned on image instructions.
Question 5In a prompt, why does it matter whether the image comes before or after the question?Think it through, then reveal
The language model is causal: a token can only attend to tokens before it. Question tokens placed after the image can look at it while being processed; question tokens placed before it cannot. The answer comes after both either way, but putting the image first lets the whole question be read in light of the picture.
Question 6Why does a speech model read a log-mel spectrogram rather than raw samples?Think it through, then reveal
Raw audio is 16,000 numbers a second with pitch hidden in the wiggles. The spectrogram makes pitch explicit (which frequencies sound at each moment), the mel bands spend detail where ears and speech need it, the log matches how loudness is heard, and 100 frames a second is far fewer positions to attend over.
Question 7How can a model that only predicts tokens produce an image or a voice?Think it through, then reveal
Encode pictures or sounds into vectors, snap each vector to its nearest entry in a learned codebook, and use the entry numbers as extra vocabulary. The model learns to predict those ids after a caption, and a decoder turns generated ids back into pixels or a waveform. The alternative is to have the model condition a diffusion model that paints the image.
Question 8A 20-minute video goes to a model with a 128,000-token window. What goes wrong, and what can be done?Think it through, then reveal
At one frame a second and 256 tokens a frame it is 1,200 × 256 = 307,200 tokens, more than twice the window, before any question. Options: sample far fewer frames (one every 5 seconds gives 61,440), merge neighbouring patch tokens, lower the resolution, transcribe the soundtrack to text, or split the video and summarise the parts.
Primary sources
The papers behind this lesson
Showed that a plain transformer over image patches matches convolutional networks when trained on enough data: the Vision Transformer.
Read the annotated companion →The paper ↗Trained an image encoder and a text encoder into one shared space; its image encoder is the starting point of many vision-language models.
Read the annotated companion →The paper ↗Connected a frozen vision encoder to a frozen language model through new cross-attention layers, handling images and video interleaved with text.
Read the annotated companion →The paper ↗The encode, project and splice recipe, trained in two stages: align the projector, then tune on image instructions.
Read the annotated companion →The paper ↗An encoder over log-mel spectrograms and a text decoder, trained on 680,000 hours of transcribed audio.
Read the annotated companion →The paper ↗Learned a codebook inside an autoencoder, turning images and audio into discrete tokens.
Read the annotated companion →The paper ↗Residual vector quantization for audio: a stack of codebooks, each encoding the previous one's error.
The paper ↗Generated images as sequences of discrete codebook tokens after the text, with one transformer.
The paper ↗Extended patches through time as tubelets, and compared ways to factorise attention over space and time.
The paper ↗Researcher's shelf
Further reading
- Dosovitskiy et al., An Image is Worth 16x16 Words (ViT, 2020): https://arxiv.org/abs/2010.11929
- Liu et al., Visual Instruction Tuning (LLaVA, 2023): https://arxiv.org/abs/2304.08485
- Liu et al., Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, 2023): https://arxiv.org/abs/2310.03744
- Li et al., BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (2023): https://arxiv.org/abs/2301.12597
- Alayrac et al., Flamingo (2022): https://arxiv.org/abs/2204.14198
- Radford et al., Whisper (2022): https://arxiv.org/abs/2212.04356 and its code: https://github.com/openai/whisper
- Défossez et al., High Fidelity Neural Audio Compression (EnCodec, 2022): https://arxiv.org/abs/2210.13438
- van den Oord et al., Neural Discrete Representation Learning (VQ-VAE, 2017): https://arxiv.org/abs/1711.00937
- Chameleon Team, Chameleon: Mixed-Modal Early-Fusion Foundation Models (2024): https://arxiv.org/abs/2405.09818
- Arnab et al., ViViT: A Video Vision Transformer (2021): https://arxiv.org/abs/2103.15691
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.