The lesson in one minute
What you'll be able to explain
- One idea: every modality becomes a sequence of vectors as wide as the language model's word vectors; after that it is ordinary attention.
- Images: a Vision Transformer cuts a picture into patches, projects each one, adds a position, and runs non-causal encoder blocks. Tokens = (H/P)·(W/P), so they grow with the square of the resolution.
- Connecting to a language model: a small projector maps image vectors into the embedding space, and they are spliced into the prompt. Train the projector first with everything frozen, then fine-tune on image instructions.
- Audio: waveform, then a log-mel spectrogram (a Fourier transform per 25 ms slice, pooled into ear-shaped bands), then an encoder: about 50 tokens a second.
- Generating pictures and sound: snap vectors to a learned codebook so they become ids a language model can predict like words; or hand off to a diffusion model.
- Video and cost: frames multiply image tokens by time; sample frames, merge tokens, and remember that text is by far the densest input.
Level 1
The practitioner's guide
In one sentence
A multimodal model turns images, audio and video into sequences of vectors the same width as a language model's word vectors, so one transformer can read (and, with a codebook, write) all of them; for a practitioner, every picture, second of sound or frame of video is a number of tokens, and that number is the cost, the latency and the limit.
When you need it
You need this lesson the moment a model has to read something that isn't text: a screenshot, a chart, a scanned form, a voice message, a meeting recording, a video clip. Whether you call a hosted vision model or build with an open one, the same questions decide the outcome: how many tokens each input becomes, at what resolution, in what order in the prompt, and whether the model needs the signal at all or only its words. You don't need a multimodal model when the words are all that matters: a transcript of speech is about 3.2 tokens a second where audio tokens are 50 a second (this lesson's rates), and a caption is a few dozen tokens where a 336 × 336 image is 576. The number that shows how fast the naive approach fails: send a 10-second clip at 30 frames a second with 256 tokens a frame and it is 76,800 tokens, more than half of a 128,000-token window, before a single word of the question. Keep one frame a second and it is 2,560.
Your options
From the cheapest to the most control:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| Turn the signal into text first | Transcribe audio (Whisper), caption or OCR an image, then use a text model | The densest input there is: hours of speech in one window | Loses tone, layout, small detail and anything the captioner missed | Your pipeline, ahead of the model |
| A hosted multimodal API | Send the image or audio in the prompt; the vendor's encoder tokenizes it | No infrastructure; a documented token count per image | Tokens per image by area (a 1000 × 1000 image is 1,296 tokens on one API); resizing above a size limit | The vendor's API |
| An open vision-language model | A ViT's patch vectors pass through a projector into the prompt (LLaVA's recipe) | Full control of resolution, tiling and prompt layout | A GPU; 196 to 576 or more tokens per image at your chosen resolution | Your server |
| A compressed-visual family | New cross-attention layers (Flamingo) or a small set of learned queries (BLIP-2's Q-Former) read the image instead of splicing it | A short text sequence however many images; 32 to 64 tokens per image | Some detail lost; a different model family to adopt | The model family you download |
| Align your own projector | Freeze a vision encoder and a language model, train the small adapter between them on captions, then tune on instructions | A model that speaks your domain's pictures, cheaply (stage 1 runs in hours) | Captioned data, then instruction data; a stage-2 fine-tune for behaviour | Your training loop |
| Discrete tokens for generation | Snap image or audio vectors to a learned codebook so the model predicts them like words | One model reads and writes every modality with next-token prediction | 1,024 ids for a 256 × 256 image; 600 ids a second for codec audio | A tokenizer plus the language model |
How to choose
Start from what the model must get out of the signal, then count the tokens.
- Only the words matter (a voicemail, a lecture, a document with plain text): transcribe or OCR to text and use a text model. It is the densest and the cheapest by a wide margin.
- Layout, tone, colour or small detail matters (a chart, a screenshot, a form, a hesitant customer): send the signal itself, at the smallest resolution that still shows the detail, cropped to the region that matters.
- Long video: sample frames (one a second for a lecture, more for anything fast), merge neighbouring patch tokens, and send the soundtrack as a transcript alongside.
- Your own domain, weak results from general models: align a projector on your captions first (cheap and safe, since both big models stay frozen), then instruction-tune.
- Generating pictures or speech from a language model: discrete codebook
tokens for simplicity, or hand the language model's output to a diffusion
model (
primer.ml.generative.diffusion) for finer detail. - Whatever you pick, put the image before the question. The language model is causal, so text placed before an image is scored as if the image were not there; the hosted API's own guidance says the same.
What it costs
Tokens grow with the square of the side: a 224 × 224 image in 16-pixel patches is 196 tokens, a 336 × 336 image in 14-pixel patches is 576, and a 1008 × 1008 image with 2 × 2 neighbours merged is 1,296. Hosted APIs bill the same way: on Claude's API each 28 × 28 pixel block is a visual token, so a 1000 × 1000 image is 1,296 tokens, larger images are downscaled to a long-edge limit (1568 pixels on the standard tier), and at $1 per million input tokens that image costs about $1.30 per thousand images. Audio is 50 encoder tokens a second on Whisper's design (a log-mel spectrogram at 100 frames a second, halved by a stride-2 convolution), so a 128,000-token window holds about 43 minutes; speech as text holds about 11 hours; video sampled at one frame a second and 256 tokens a frame holds about 8 minutes, and full-rate video under 20 seconds. Every one of those tokens lengthens the prefill before the first word of the answer and competes with instructions and retrieved documents for attention. Training cost is lopsided: stage-1 alignment trains a projector of a few million numbers between models of billions (this lesson's toy takes 80 pictures from 24% to 99% correct), and BLIP-2 beat the 80-billion-parameter Flamingo on zero-shot VQAv2 by 8.7% with 54 times fewer trainable parameters; it is stage 2, the instruction data, that decides how the model behaves.
What breaks
- Objects that aren't there. A weakly aligned model describes what pictures like this usually contain; the hosted API warns of the same with low-quality, rotated or very small images (under 200 pixels). Send clear images at a usable size, and verify anything that matters.
- The image after the question. The question's tokens cannot attend forward to an image that follows them. Image first, then the question.
- Unreadable small text. A screenshot downscaled to the size limit loses its small print. Crop to the region, or tile the page, rather than sending one huge image.
- The context window full of frames. 20 minutes of video at one frame a second and 256 tokens a frame is 307,200 tokens, over twice a 128k window. Sample less often, merge tokens, transcribe the audio, or split the clip.
- Sampling that misses the moment. One frame a second is fine for a lecture and useless for a golf swing. Match the frame rate to the speed of what you are looking for.
- Transcribing away the signal. A transcript drops hesitation, tone and who spoke; a caption drops layout. Use text only when the words are enough.
- A codebook that fits nothing. A random codebook is useless at every size (error 0.6 to 2.1 against 0.09 to 0.32 for a learned one in this lesson's sweep); codebooks are learned on the data, and a tokenizer from one domain misbehaves on another.
In the wild
The Vision Transformer (Dosovitskiy et al.) is the front end: a pure transformer over 16 × 16 patches that matched convolutional networks given enough data, and CLIP's ViT (Radford et al.) is the one most vision-language models start from. LLaVA (Liu et al.) is the encode, project and splice recipe, trained on instruction data generated with a language model, scoring 85.1% of GPT-4 on its own multimodal benchmark; Flamingo (Alayrac et al.) reads images through new cross-attention layers between frozen models and learns tasks from a few examples in the prompt; BLIP-2 (Li et al.) compresses each image through a Querying Transformer. Whisper (Radford et al.) is the audio side: an encoder over log-mel spectrograms and a text decoder trained on 680,000 hours of audio, used zero-shot for transcription. For generation, VQ-VAE and DALL-E turned images into codebook ids for a transformer, SoundStream and EnCodec do it for audio with stacked residual codebooks, and Chameleon trains one model on interleaved text and image tokens from the start. Hosted vision APIs expose the same arithmetic as a price list: Claude's vision documentation, the source of the numbers above, counts each 28 × 28 pixel block as a visual token and caps images by long edge and token count. Every paper is linked at the end of the lesson.
Go deeper
Level 2 builds each front end by hand: a 4 × 4 picture cut into four tokens, a Vision Transformer with its projector spliced into a tiny language model and aligned on 80 pictures, a Fourier transform and a log-mel spectrogram from eight samples upward, a codebook learned by k-means with residual stages, and the token arithmetic for video and the context budget. If you only needed to count the tokens and choose, you are done.
Level 2
How it works, from scratch
Imagine a brilliant reader who can take in only one thing: a long row of
index cards, each card holding a short list of numbers. That is a language
model. Every word it reads arrives as one card, the word's vector (see
primer.ml.tokenization and primer.ml.transformer). It has never seen a
photograph or heard a voice.
To tell this reader about a photo, you cut the photo into small squares and write one card per square. To tell it about a voice recording, you write one card for every fiftieth of a second of sound. If the cards are written in the same "handwriting" as the word cards, the reader handles a photo the way it handles a sentence: it pays attention across all the cards at once.
That is the whole idea of a multimodal model, a model that takes in more than one modality (kind of input: text, images, audio, video):
Turn every modality into a sequence of vectors ("tokens") that a transformer can attend over.
Everything in this lesson is a way of making those cards: for images, for sound, for video, and, run backwards, for a model that produces pictures and speech.
Figure 1 · Diagram
flowchart LR IMG["Image"] --> VE["Vision encoder<br/>(patches to vectors)"] --> P1["Projector"] AUD["Audio"] --> SP["Spectrogram"] --> AE["Audio encoder"] --> P2["Projector"] TXT["Text"] --> TOK["Tokenizer"] --> EMB["Embedding table"] P1 --> SEQ["One sequence of vectors,<br/>all the same width"] P2 --> SEQ EMB --> SEQ SEQ --> LM["Decoder language model"] --> OUT["Next token:<br/>a word, or a codebook id<br/>for a picture or a sound"]
Chapter 1
A tiny worked example: a 4×4 picture becomes 4 tokens
Take a grayscale picture of 4×4 pixels, each pixel a brightness from 0 to 15:
0 1 | 2 3
4 5 | 6 7
------+------
8 9 | 10 11
12 13 | 14 15
- Cut it into 2×2 patches (the lines above): 4 patches.
- Flatten each patch into a list, reading row by row: the top-right patch becomes (2, 3, 6, 7).
- Project each list to the model's width, here 2 numbers, by multiplying by a matrix whose first column adds up the patch's top row and whose second column adds up its bottom row.
- Add a position: the patch's (row, column) in the grid, so the model knows where each patch came from.
| Patch | Pixels | Projected | + position | Token |
|---|---|---|---|---|
| top left | 0, 1, 4, 5 | (1, 9) | (0, 0) | (1, 9) |
| top right | 2, 3, 6, 7 | (5, 13) | (0, 1) | (5, 14) |
| bottom left | 8, 9, 12, 13 | (17, 25) | (1, 0) | (18, 25) |
| bottom right | 10, 11, 14, 15 | (21, 29) | (1, 1) | (22, 30) |
The picture is now four tokens of two numbers each: exactly the shape of input a transformer reads. A real model does the same with bigger numbers: patches of 14 or 16 pixels, and hundreds or thousands of numbers per token.
In code: worked_example_patches runs these four steps on this picture and returns the patches, the projections and the tokens.
Chapter 2
Images: a Vision Transformer from scratch
primer.ml.cnn_rnn introduced the idea of patches as tokens. Here we build
the whole front end, the Vision Transformer (ViT), and count what it
costs.
Cutting and counting
Everyday picture Lay a sheet of graph paper over a photo and cut along every fourth line. You get a grid of small tiles, and you can hand them over one by one, left to right and top to bottom, like the words of a sentence.
Tiny example A 224×224 image in 16-pixel patches is a grid of 224 / 16 = 14 patches down and 14 across: 14 × 14 = 196 tokens. A 336×336 image in 14-pixel patches is 24 × 24 = 576 tokens.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the image's height and width, in pixels | 224, 224 | |
| the side of one square patch, in pixels | 16 | |
| how many patches fit down the image (must be a whole number, so images are resized first) | 14 | |
| how many patches fit across | 14 | |
| multiply | ||
| the number of tokens the image becomes | 196 |
Each token starts life as numbers, where is the number of colour channels (3 for red, green, blue): 16 · 16 · 3 = 768.
In words: "the number of image tokens is the number of patches down times the number of patches across."
With the numbers: (224 / 16) · (224 / 16) = 14 · 14 = 196; (336 / 14) · (336 / 14) = 24 · 24 = 576; the worked 4×4 picture in 2-pixel patches gives 2 · 2 = 4.
In Python:
H, W, P = 224, 224, 16
# H/P patches down, times W/P across
(H // P) * (W // P) # → 196
H, W, P = 336, 336, 14
(H // P) * (W // P) # → 576
# the worked 4×4 picture in 2-pixel patches
(4 // 2) * (4 // 2) # → 4
Figure 2 · Drawn from the lesson's code
A 32 by 32 picture of a sun over a striped field, cut into 16 numbered 8-pixel patches, and the 16-row matrix those patches become
Figure 3 · Drawn from the lesson's code
Tokens per image rise with the square of the side length: 576 at 336 pixels, over 9,000 at 1,344 pixels with 14-pixel patches, and a quarter of that when 2 by 2 neighbours are merged
In code: image_to_patches cuts and flattens an image (the cutting is primer.ml.cnn_rnn.patchify), refusing sizes the patch does not divide, and count_image_tokens is the formula above, with an optional merge.
Why it matters in practice. Resolution is a cost dial. Small text in a
screenshot needs a high resolution to be legible, and every doubling of the
side costs four times the tokens (and attention cost grows faster still; see
primer.ml.attention). Systems resize images to a supported size, cut very
large ones into tiles, and merge neighbouring patches to keep the count
manageable.
From patch to token: projection and position
Everyday picture Every tile gets the same questionnaire: "how bright is your top half? your bottom half? is there an edge?" The answers become the tile's card. Then a sticker with the tile's grid address goes on the card, so shuffling the cards loses nothing.
Tiny example In the worked example, the questionnaire had two questions (top-row total, bottom-row total), and the sticker was the patch's (row, column).
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape | In the example (top-right patch) |
|---|---|---|---|
| which patch, counting in reading order from 0 | 1 | ||
| patch flattened into a row of pixels | (2, 3, 6, 7) | ||
| the learned patch-embedding matrix: one column per output number (a bias vector is usually added too; it is left out here) | the 4 × 2 matrix below | ||
| a matrix multiply: each output number is the dot product of the patch with one column of | (5, 13) | ||
| the learned position vector for slot | (0, 1) | ||
| the finished token for patch | (5, 14) | ||
| the model width: how many numbers per token | 2 |
in the example is the rows (1, 0), (1, 0), (0, 1), (0, 1): the first two pixels (the patch's top row) feed output 1, the last two (its bottom row) feed output 2.
In words: "each token is its patch multiplied by a learned matrix, plus a learned vector that marks where the patch sits."
With the numbers: (2, 3, 6, 7) · = (2 + 3, 6 + 7) = (5, 13); adding the position (0, 1) gives (5, 14).
In Python:
x_i = [2, 3, 6, 7]
W_E = [[1, 0], [1, 0], [0, 1], [0, 1]]
p_i = [0, 1]
# x_i W_E: dot the patch with each column of W_E
xW = [sum(x * W_E[r][c] for r, x in enumerate(x_i)) for c in range(2)]
xW # → [5, 13]
# + p_i
[a + b for a, b in zip(xW, p_i)] # → [5, 14]
Why the position vector? Attention compares every token with every other
and ignores their order (see primer.ml.positional). Without , a sky
patch at the top and the same sky colour in a puddle at the bottom would be
the same token, and "the sun is above the field" could not be expressed.
Figure 4 · Diagram
flowchart LR IMG["Image<br/>H × W × C"] --> CUT["Cut into N patches<br/>N × P·P·C"] CUT --> PROJ["Multiply by W_E<br/>N × d"] PROJ --> POS["Add position vectors<br/>N × d"] POS --> ENC["Encoder blocks<br/>every patch attends to every patch<br/>(no causal mask)"] ENC --> OUT["N image vectors<br/>N × d"]
primer.ml.transformer, run with the causal mask switched off. A sentence
is read left to right, so a text decoder hides the future; a picture has no
future, so every patch may look at every other, above, below and to either
side. The output has one vector per patch, now informed by the whole image.In code: VisionEncoder holds W_E, the position table and a stack of primer.ml.transformer.TransformerBlock with causal=False; VisionEncoder.embed is the formula above, and calling the encoder runs the blocks.
Why it matters in practice. The vision encoder inside most
vision-language models is a ViT that was first trained as the image half of
CLIP (see primer.ml.embeddings.contrastive). CLIP training pulled its
image vectors towards the vectors of matching captions, so its outputs
already carry meaning a language model can use. That is why builders start
from it instead of training vision from scratch.
Chapter 3
Connecting vision to a language model
The projector: a plug adapter
Everyday picture Your laptop charger has the right voltage but the wrong plug for the wall socket abroad. A small adapter changes the shape, not the power. The vision encoder speaks in vectors of its own width and style; the language model expects vectors shaped like its word embeddings. A projector is the adapter between them.
Tiny example A vision vector with 2 numbers, (1, 2), must become a language-model vector with 3 numbers.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | Shape | In the example |
|---|---|---|---|
| one image token from the vision encoder | (1, 2) | ||
| the projector's learned matrix | rows (1, 0, 1) and (0, 1, 1) | ||
| the same token, now shaped like a word embedding | (1, 2, 3) | ||
| the widths of the two models | 2 and 3 |
In words: "multiply each image vector by one learned matrix to turn it into a vector the language model can read."
With the numbers: (1, 2) · = (1·1 + 2·0, 1·0 + 2·1, 1·1 + 2·1) = (1, 2, 3).
Level 3: in Python
v = [1, 2]
W_P = [[1, 0, 1], [0, 1, 1]]
# v W_P: dot v with each column of W_P
[sum(v[r] * W_P[r][c] for r in range(2)) for c in range(3)] # → [1, 2, 3]
Many models use two such layers with a GELU between them (a smooth
"keep the positives" function; see primer.ml.transformer), which is a
small feed-forward network. Either way, the projector is tiny next to the
two models it joins: a few million numbers between models with billions.
In code: Projector is a single linear layer by default and a two-layer MLP when given a hidden width; worked_example_projection computes the (1, 2, 3) above.
Splicing image tokens into the prompt
Everyday picture Writing a letter and taping a strip of photos into the middle of a sentence. The reader reads the words, then the photos, then the rest of the words, in order.
Tiny example The prompt "what is this <image>" has four text tokens,
one of them a placeholder. The placeholder is replaced by the image's 4
projected vectors, so the language model reads 3 + 4 = 7 rows.
| Step | Shape (toy model) | Meaning |
|---|---|---|
| image | (8, 8) | an 8×8 grayscale picture |
| patches | (4, 16) | four 4×4 patches, flattened |
| vision encoder output | (4, 16) | four image vectors of width 16 |
| projector output | (4, 24) | four image tokens, as wide as the language model |
| text token vectors | (3, 24) | "what", "is", "this" from the embedding table |
| spliced sequence | (7, 24) | text, then image, in prompt order |
| logits | (7, 12) | a score for each of the 12 vocabulary words, at every position |
Figure 6 · Diagram
flowchart LR IMG["Image"] --> VE["Vision encoder<br/>4 × 16"] --> PR["Projector<br/>4 × 24"] TXT["Prompt: what is this,<br/>then the image placeholder"] --> TE["Look up text tokens<br/>3 × 24"] PR --> SPL["Splice at the placeholder<br/>7 × 24"] TE --> SPL SPL --> POS["+ positions"] --> DEC["Decoder blocks<br/>(causal)"] --> LOG["Logits<br/>7 × 12"]
Because the decoder is causal, a token can use the image only if it comes after the image. Text before the picture is scored exactly as if there were no picture. That is why prompts usually put the image first and the question after it.
In code: VisionLanguageModel.splice builds the spliced sequence and a flag for each row saying whether it is an image token; calling VisionLanguageModel runs primer.ml.transformer.TinyGPT's blocks on it; build_toy_vlm wires up the toy model in the table.
Why it matters in practice. This "encode, project, splice" design (as in LLaVA) is the simplest and most common. Two other families exist. One adds new cross-attention layers inside the language model that look at the image vectors (as in Flamingo), leaving the text sequence short. Another first compresses any image into a fixed small number of tokens, such as 32 or 64, with a little attention module of learned queries (as in BLIP-2's Q-Former). Both trade some detail for fewer tokens.
How such a model is trained
Everyday picture Two experts who don't share a language: a photographer and a writer. You hire an interpreter. First, the interpreter learns vocabulary while both experts carry on exactly as they are: the photographer points at pictures, the writer names them. Then all three practise real conversations together, and the writer is allowed to adapt a little too.
That is the usual two-stage recipe:
- Alignment. Freeze the vision encoder and the language model. Train only the projector on image-caption pairs, so the projected image tokens make the language model produce the caption.
- Visual instruction tuning. Train the projector and the language model
(fully, or with LoRA; see
primer.ml.training_stages) on images paired with instructions and answers: "what is unusual about this picture?" followed by a good answer.
In both stages the loss is ordinary next-token cross-entropy (the
negative log of the probability given to the right token; see
primer.ml.losses), counted only on the answer tokens. The model is not
graded on predicting the image tokens or the question it was given.
Tiny example The answer is "horizontal stripes", two tokens. The model gives the right first token probability 0.5 and the right second token 0.25.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the positions of the answer tokens (the loss mask keeps only these) | 2 positions | |
| how many answer tokens there are | 2 | |
| "for each position in the answer" | ||
| the correct token at position | "horizontal", then "stripes" | |
| every token before position | ||
| the probability the model gives the correct token, given () the image and what came before | 0.5, 0.25 | |
the natural logarithm; is 0 when and grows as shrinks (see primer.notation) |
||
| the loss: the average surprise over the answer | 1.040 |
In words: "the loss is the average, over the answer tokens only, of how surprised the model was by each correct token."
With the numbers: −(log 0.5 + log 0.25) / 2 = (0.693 + 1.386) / 2 = 1.040.
Level 3: in Python
import math
# probabilities of the right answer tokens
p = [0.5, 0.25]
# −(1/|A|) Σ log p
round(-sum(math.log(q) for q in p) / len(p), 3) # → 1.04
Figure 7 · Diagram
flowchart TB
subgraph S1["Stage 1: alignment (image-caption pairs)"]
direction LR
V1["Vision encoder<br/>frozen"] --> P1["Projector<br/>TRAINED"] --> L1["Language model<br/>frozen"]
end
subgraph S2["Stage 2: visual instruction tuning (image, question, answer)"]
direction LR
V2["Vision encoder<br/>frozen"] --> P2["Projector<br/>trained"] --> L2["Language model<br/>trained or LoRA"]
end
S1 --> S2
Our toy does stage 1. Its vision encoder and language model are random and frozen; only the projector learns, from 80 small pictures of four patterns (horizontal, vertical, diagonal, checkered), to make the language model's output layer pick the right pattern word.
Figure 5 · Drawn from the lesson's code
Stage-1 alignment: the caption loss falls from 2.41 to 0.33, and on 80 new images the right word ranks first 99% of the time, up from 24%
In code: align_projector trains only the projector's weights with cross-entropy on the caption word and returns the loss and accuracy per step; caption_accuracy measures how often the right word wins. The toy scores a pooled image vector directly against the language model's token table (the last step of the real path) rather than back-propagating through every layer.
Why it matters in practice. Stage 1 is cheap enough to run in hours, because almost every weight is frozen. Stage 2 decides the model's behaviour: the quality and variety of the image-instruction data matter more than its size. A common failure of a weakly aligned model is describing objects that are not in the picture, a visual form of hallucination.
Chapter 4
Audio: from a waveform to tokens
Sound is a list of numbers
Everyday picture A microphone is a tiny eardrum. It measures air pressure many thousands of times a second, and the recording is just that list of measurements: the waveform. Drawn on paper, it looks like a seismograph trace.
Tiny example Speech is usually recorded at a sample rate of 16,000 measurements per second, so one second is 16,000 numbers and a 30-second clip is 480,000. Used directly as tokens that would be ruinous, and the thing that matters, which pitches are sounding, is hidden in the wiggles (top panel of the spectrogram figure below).
How much of each pitch: the Fourier transform
Everyday picture A prism splits white light into its colours. The Fourier transform splits a sound into its pitches. It works by holding a pure wave of each pitch up against the recording and asking "how well do you line up?". A pitch that is present lines up again and again and scores high; a pitch that is absent lines up as often as it clashes and scores zero.
Tiny example Eight samples of a wave that goes up and down twice: (1, 0, −1, 0, 1, 0, −1, 0). Line it up against a wave that also cycles twice, cos: (1, 0, −1, 0, 1, 0, −1, 0). Multiply matching positions and add: 1 + 0 + 1 + 0 + 1 + 0 + 1 + 0 = 4. A wave that cycles once agrees for half its length and disagrees for the other half: it scores 0.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the -th sample of the recording | (1, 0, −1, 0, 1, 0, −1, 0) | |
| how many samples | 8 | |
| a counter over the samples, from 0 to | 0, 1, …, 7 | |
| which pitch we are testing: a wave that completes cycles in the samples | 2 | |
| the cosine and sine waves (sine is the same wave shifted a quarter cycle, so a pitch that starts at a different moment is still caught) | ||
| one full cycle, in the units cos and sin use | ||
| add up over every sample | ||
| the length of the pair (cosine score, sine score) | ||
| how much of pitch the recording holds | 4 |
Bin is the frequency in hertz (cycles per second), where is the sample rate. If the 8 samples were taken over one second, and bin 2 is 2 Hz. Textbooks write the same thing with complex numbers, ; the cosine part and the sine part are exactly the two sums above.
In words: "for each pitch, correlate the recording with a cosine and a sine of that pitch, and take the length of the two scores."
With the numbers: for the cosine sum is 4 and the sine sum is 0, so . For both sums are 0. Bin 6 also scores 4: with only 8 samples, a wave cycling 6 times is indistinguishable from one cycling 8 − 6 = 2 times, so the top half of the bins mirrors the bottom half and only bins 0 to are kept.
In Python:
import math
x = [1, 0, -1, 0, 1, 0, -1, 0]
N = len(x)
def magnitude(k):
c = sum(x_n * math.cos(2 * math.pi * k * n / N) for n, x_n in enumerate(x))
s = sum(x_n * math.sin(2 * math.pi * k * n / N) for n, x_n in enumerate(x))
return math.sqrt(c ** 2 + s ** 2)
[round(magnitude(k), 6) for k in range(N)] # → [0.0, 0.0, 4.0, 0.0, 0.0, 0.0, 4.0, 0.0]
# f_k = k · f_s / N, with 8 samples a second
2 * 8 / N # → 2.0
In code: dft_magnitudes is this formula for every k at once, written from the definition (the tests check it against NumPy's fast Fourier transform, which computes the same numbers far faster).
The spectrogram: one Fourier transform per slice
Everyday picture A piano roll, or sheet music: time runs left to right, pitch runs bottom to top, and a mark means "this note is sounding now". To make one from a recording, listen through a short window, say 25 milliseconds, ask "which pitches are in here?", slide the window along a little, and ask again. That is the short-time Fourier transform (STFT), and its picture is a spectrogram.
Each slice is faded in and out with a Hann window (a smooth hump from 0 up to 1 and back) before measuring, because chopping a wave off abruptly creates a click, and a click contains every pitch at once.
Tiny example Speech systems commonly use 25 ms windows (400 samples at 16 kHz) that start every 10 ms (160 samples). How many windows fit in one second?
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the recording's length, in samples | 16,000 | |
| the window length, in samples | 400 (25 ms) | |
| the hop: how far the window moves each step | 160 (10 ms) | |
| floor: round down to a whole number (a part-window at the end is dropped) | ||
| the number of frames (columns of the spectrogram) | 98 |
In words: "one window fits at the start; after that, count how many whole hops still leave room for a full window."
With the numbers: 1 + ⌊(16,000 − 400) / 160⌋ = 1 + ⌊97.5⌋ = 98 frames per second, which is why speech models talk about "100 frames a second".
Level 3: in Python
L, N, H = 16000, 400, 160
# 1 + ⌊(L − N) / H⌋
1 + (L - N) // H # → 98
Figure 10 · Diagram
flowchart LR W["Waveform<br/>L samples"] --> F["Slice into frames<br/>T × N"] F --> HW["Fade each frame<br/>(Hann window)"] HW --> DFT["Fourier transform<br/>per frame<br/>T × (N/2 + 1)"] DFT --> MEL["Pool into mel bands<br/>T × n_mels"] MEL --> LOG["Take the log<br/>log-mel spectrogram"]
Figure 8 · Drawn from the lesson's code
A 20-millisecond waveform of wiggles, then a spectrogram of the same second of sound showing a flat line at 1000 Hz and a line rising from 200 to 3000 Hz, then the log-mel version where the rising line curves
In code: stft slices, windows (with hann) and measures every frame; frame_count is the formula above; log_mel_spectrogram adds the mel pooling and the log.
The mel scale: spacing pitch the way ears do
Everyday picture On a piano, every octave takes up the same width of keyboard, yet each octave doubles the frequency: the A keys are at 110, 220, 440 and 880 Hz. Ears work the same way above about 1,000 Hz: a jump from 1,000 to 2,000 Hz sounds about as big as a jump from 2,000 to 4,000. The mel scale relabels frequencies so that equal steps in mel sound like equal steps in pitch.
Tiny example By construction 1,000 Hz is 1,000 mel. The 7,000 Hz from 1,000 to 8,000 Hz shrinks to just 1,840 mel.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| a frequency in hertz | 700 | |
| frequency measured in units of 700 Hz; below about 700 Hz the scale is nearly straight, above it the log takes over | 1 | |
the base-10 logarithm: "10 to what power gives this?"; it turns ratios into equal steps (see primer.notation) |
||
| 2595 | a constant chosen so that 1,000 Hz comes out as 1,000 mel | |
| the same frequency in mel | 781.2 |
In words: "mel is a log of the frequency, gently straightened below 700 Hz and scaled so that 1,000 Hz is 1,000 mel."
With the numbers: 2595 · log₁₀(1 + 700/700) = 2595 · 0.301 = 781.2; 1,000 Hz gives 1,000.0; 4,000 Hz gives 2,146.1; 8,000 Hz gives 2,840.0.
Level 3: in Python
import math
def mel(f):
return 2595 * math.log10(1 + f / 700)
round(mel(700), 1) # → 781.2
round(mel(1000), 1) # → 1000.0
round(mel(4000), 1) # → 2146.1
A mel filterbank turns the hundreds of frequency bins into a few dozen bands: triangles evenly spaced in mel, so narrow at low frequencies and wide at high ones. Each band adds up the energy under its triangle. Loudness is also heard by ratio, so the last step takes the log.
Figure 9 · Drawn from the lesson's code
Left: the mel curve rises steeply to 1000 mel at 1000 Hz and then flattens, reaching only 2840 at 8000 Hz. Right: ten triangular filters, narrow below 1000 Hz and ever wider above
In code: hz_to_mel is the formula, and mel_filterbank builds the triangles, evenly spaced in mel.
Speech recognition: an encoder over frames, a decoder for text
Everyday picture A court stenographer listens to a whole sentence, then types it out word by word, glancing back at what they heard as they go.
Tiny example Take a 30-second clip. With 10 ms hops it is 3,000 frames of 80 mel bands each. A small convolution that moves two frames at a time halves that to 1,500 encoder positions: 50 audio tokens per second. A transformer encoder attends over all 1,500 at once, and a decoder writes the transcript as text tokens, attending both to its own words so far and to the encoder's output. This is the design of Whisper.
Figure 11 · Diagram
flowchart LR A["30 s of audio"] --> LM["Log-mel spectrogram<br/>3000 frames × 80 bands"] LM --> CV["Convolution, stride 2<br/>1500 positions"] CV --> ENC["Encoder blocks<br/>(no causal mask)"] ENC --> X["Cross-attention:<br/>decoder reads the audio"] DEC["Decoder blocks<br/>(causal)"] --> X X --> TXT["Text tokens,<br/>one at a time"] TXT -.-> DEC
primer.ml.transformer for encoders and decoders). The dotted arrow is the
generation loop: each new word is fed back in.In code: audio_tokens counts encoder positions from a clip's length, the hop and the downsampling.
Why it matters in practice. The same audio encoder can feed a general language model instead of a dedicated decoder: add a projector, splice the audio tokens into the prompt, and the model can answer questions about a recording, exactly as with images. Fifty tokens a second is manageable for a short voice message and expensive for an hour-long meeting.
Chapter 5
Generating images and speech: discrete tokens from a codebook
So far the model reads pictures and sound. To write them, a language model needs them as something it can predict one at a time from a fixed vocabulary, the way it predicts words.
Everyday picture A paint-by-numbers kit comes with a palette of numbered pots. Any small area of a picture is described by the number of the pot closest to its colour. The whole picture becomes a grid of pot numbers, and anyone with the same palette can repaint it. The palette is a codebook; snapping to the nearest entry is vector quantization.
Tiny example A codebook of three 2-number entries: , , . The vector has squared distances 0.85, 0.05 and 1.45 to them, so it becomes id 1. Decoding id 1 gives back (1, 0): close, not exact.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| one vector to encode: an image patch's vector or a slice of sound | (0.9, 0.2) | |
| codebook entry number | ||
| the codebook size: how many entries | 3 | |
| " ranges over the entry numbers" | 0, 1, 2 | |
| squared distance: subtract, square each difference, add | ||
| "the that gives the smallest value" (the position of the minimum, not the minimum itself) | 1 | |
| the id that replaces | 1 |
In words: "replace each vector by the number of its nearest codebook entry."
With the numbers: distances squared to are 0.81 + 0.04 = 0.85, 0.01 + 0.04 = 0.05 and 0.81 + 0.64 = 1.45; the smallest is 0.05, at .
Level 3: in Python
z = [0.9, 0.2]
codebook = [[0, 0], [1, 0], [0, 1]]
# ‖z − c_k‖² for each entry
d2 = [sum((a - b) ** 2 for a, b in zip(z, c)) for c in codebook]
[round(d, 2) for d in d2] # → [0.85, 0.05, 1.45]
# arg min: the position of the smallest
d2.index(min(d2)) # → 1
Figure 13 · Diagram
flowchart LR IN["Picture or sound"] --> E["Encoder<br/>vectors"] E --> Q["Snap to nearest<br/>codebook entry"] Q --> IDS["Ids, like word tokens<br/>e.g. 17, 803, 4, ..."] IDS --> LM["Language model<br/>learns to predict ids<br/>after text"] LM --> GEN["Generated ids"] GEN --> LOOK["Look up codebook<br/>vectors"] LOOK --> D["Decoder"] --> OUTP["Pixels or waveform"]
primer.ml.generative.autoencoders).How is the codebook chosen? It is learned so that the entries sit where the
data actually is, which amounts to k-means clustering (see
primer.ml.embeddings.clustering): snap every vector to its nearest entry,
move each entry to the average of the vectors that chose it, and repeat.
For audio, one codebook is rarely precise enough, so neural audio codecs
use residual vector quantization: a second codebook encodes what the
first one missed, a third what the second missed, and so on. Each moment
of sound becomes a small stack of ids.
Figure 12 · Drawn from the lesson's code
On new patches, a learned codebook's error falls from 0.32 to about 0.09 as it grows to 64 entries while random codebooks stay far above it; stacking residual codebooks of 8 entries keeps pushing the error down
In code: quantize is the arg min above and dequantize the lookup; kmeans_codebook learns a codebook; learn_residual_codebooks and residual_quantize stack codebooks on the leftovers.
Why it matters in practice. Discrete tokens let one model read and
write every modality with one mechanism, next-token prediction. The costs
are real: a 256×256 image compressed eight-fold per side is 32 × 32 =
1,024 ids (the scheme of the original DALL·E), and a neural audio codec
running at 75 frames a second with 8 codebooks writes 600 ids a second. The
other road to generating images is diffusion (see
primer.ml.generative.diffusion): the language model supplies text or
vectors that condition a diffusion model, which paints the pixels in many
small denoising steps. Discrete tokens are simpler to bolt onto a language
model; diffusion usually renders finer detail.
Chapter 6
Video: pictures over time
Everyday picture A flipbook. Each page is a picture, and neighbouring pages are almost identical. To follow the story you don't need to look at every page; a glance at every tenth page shows what happens.
Tiny example A 10-second clip at 30 frames a second, with each frame cut into 256 tokens (224×224 in 14-pixel patches): 300 frames × 256 = 76,800 tokens, more than half of a 128,000-token context, for ten seconds. Keeping one frame a second gives 10 × 256 = 2,560.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the clip's duration, in seconds | 10 | |
| frames kept per second (the sampling rate, often far below the video's own frame rate) | 30, then 1 | |
| the tubelet depth: how many consecutive frames share one token, when a patch is cut through time as well as space | 1 | |
| how many frame slots become tokens | 300, then 10 | |
| tokens per frame, from the image formula | 16 · 16 = 256 | |
| the clip's total tokens | 76,800, then 2,560 |
In words: "count the frames you keep, divide by how many frames each token spans, and multiply by the tokens per frame."
With the numbers: (10 · 30 / 1) · 16 · 16 = 300 · 256 = 76,800; at one frame a second, (10 · 1 / 1) · 256 = 2,560; sampling 2 frames a second in tubelets 2 frames deep gives (10 · 2 / 2) · 256 = 2,560 again, while covering twice as many moments.
Level 3: in Python
H = W = 224
P = 14
tokens_per_frame = (H // P) * (W // P)
tokens_per_frame # → 256
D, r, t = 10, 30, 1
D * r // t * tokens_per_frame # → 76800
# one frame a second
D, r, t = 10, 1, 1
D * r // t * tokens_per_frame # → 2560
Figure 14 · Diagram
flowchart LR V["Video<br/>30 frames a second"] --> S["Sample frames<br/>e.g. 1 a second"] S --> ENC["Vision encoder per frame<br/>(or tubelets across frames)"] ENC --> M["Merge neighbouring<br/>patch tokens"] M --> PR["Projector"] PR --> SEQ["Frame tokens,<br/>with timestamps as text"] SEQ --> LM["Language model"]
In code: video_tokens is the formula, and sample_frames picks the evenly spaced frames to keep.
Why it matters in practice. Sampling is a bet that nothing important happens between the frames you keep: one frame a second is fine for a lecture and useless for a golf swing. Real systems adapt: more frames for short or fast clips, fewer and smaller ones for long recordings, and a transcript of the soundtrack alongside.
Chapter 7
What it costs: tokens per second and the context budget
Everyday picture Packing a suitcase. Text is neatly folded shirts; images are shoes; video is an inflatable boat. The suitcase (the context window) is the same size whatever you pack.
Tiny example A 128,000-token window holds 128,000 / 256 = 500 seconds, about 8 minutes, of video sampled at one frame a second, before a single word of the question. The same window holds about 11 hours of speech written out as text.
| Input | Tokens per second | Minutes in 128k tokens |
|---|---|---|
| speech written as text (about 150 words a minute, about 1.3 tokens a word) | 3.2 | 656 |
| speech as audio-encoder tokens | 50 | 43 |
| video, 1 frame a second, 256 tokens a frame | 256 | 8.3 |
| audio as codec tokens (75 frames × 8 codebooks) | 600 | 3.6 |
| video, 30 frames a second, 256 tokens a frame | 7,680 | 0.3 |
Figure 15 · Drawn from the lesson's code
On log axes, five straight lines of tokens against recording length: speech-as-text crosses a 128k window only after hours, 30-frames-a-second video within seconds
In code: seconds_that_fit divides a window by a rate, and TOKENS_PER_SECOND holds the rates in the table.
Why it matters in practice. Every image and second of audio is billed
and computed like text: it takes room in the context window, lengthens the
prefill before the first word of the answer (see primer.ml.inference),
and competes with instructions and retrieved documents for attention (see
primer.agents.context). The practical levers follow from the formulas:
send the smallest resolution that still shows the detail you need, crop to
the region that matters, sample fewer frames, and transcribe audio to text
when the words matter and the tone does not.
Test yourself
8 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1How can a chatbot "see" a photo, explained without jargon?Think it through, then reveal
The photo is cut into a grid of small squares, and each square is described by a list of numbers in the same format the model uses for words. The model then reads the squares and the words of your question together, paying attention to whichever squares help answer it. It never sees the picture as a picture: it reads it as a few hundred extra "words".
Question 2How many tokens is a 448×448 image with 14-pixel patches? And if each 2×2 block of neighbours is merged?Think it through, then reveal
448 / 14 = 32 patches a side, so 32 × 32 = 1,024 tokens. Merging 2×2 blocks divides by 4: 256 tokens. Doubling the side from 224 would have quadrupled the count.
Question 3Why does a Vision Transformer need position vectors, and why is its attention not causal?Think it through, then reveal
Attention ignores order, so without positions the model sees an unordered bag of patches and cannot tell the sky from the ground. It is not causal because a picture has no "future": every patch may use every other, and hiding half the image would only throw information away.
Question 4What does the projector do, and why train it first with both big models frozen?Think it through, then reveal
It maps each image vector into the language model's embedding space, so image tokens look like something the language model can use. Training it alone first is cheap (it is tiny) and safe (the frozen models keep all they know). Only once the two sides understand each other is the language model tuned on image instructions.
Question 5In a prompt, why does it matter whether the image comes before or after the question?Think it through, then reveal
The language model is causal: a token can only attend to tokens before it. Question tokens placed after the image can look at it while being processed; question tokens placed before it cannot. The answer comes after both either way, but putting the image first lets the whole question be read in light of the picture.
Question 6Why does a speech model read a log-mel spectrogram rather than raw samples?Think it through, then reveal
Raw audio is 16,000 numbers a second with pitch hidden in the wiggles. The spectrogram makes pitch explicit (which frequencies sound at each moment), the mel bands spend detail where ears and speech need it, the log matches how loudness is heard, and 100 frames a second is far fewer positions to attend over.
Question 7How can a model that only predicts tokens produce an image or a voice?Think it through, then reveal
Encode pictures or sounds into vectors, snap each vector to its nearest entry in a learned codebook, and use the entry numbers as extra vocabulary. The model learns to predict those ids after a caption, and a decoder turns generated ids back into pixels or a waveform. The alternative is to have the model condition a diffusion model that paints the image.
Question 8A 20-minute video goes to a model with a 128,000-token window. What goes wrong, and what can be done?Think it through, then reveal
At one frame a second and 256 tokens a frame it is 1,200 × 256 = 307,200 tokens, more than twice the window, before any question. Options: sample far fewer frames (one every 5 seconds gives 61,440), merge neighbouring patch tokens, lower the resolution, transcribe the soundtrack to text, or split the video and summarise the parts.
Primary sources
The papers behind this lesson
Showed that a plain transformer over image patches matches convolutional networks when trained on enough data: the Vision Transformer.
Read the annotated companion →The paper ↗Trained an image encoder and a text encoder into one shared space; its image encoder is the starting point of many vision-language models.
Read the annotated companion →The paper ↗Connected a frozen vision encoder to a frozen language model through new cross-attention layers, handling images and video interleaved with text.
Read the annotated companion →The paper ↗The encode, project and splice recipe, trained in two stages: align the projector, then tune on image instructions.
Read the annotated companion →The paper ↗An encoder over log-mel spectrograms and a text decoder, trained on 680,000 hours of transcribed audio.
Read the annotated companion →The paper ↗Learned a codebook inside an autoencoder, turning images and audio into discrete tokens.
Read the annotated companion →The paper ↗Residual vector quantization for audio: a stack of codebooks, each encoding the previous one's error.
The paper ↗Generated images as sequences of discrete codebook tokens after the text, with one transformer.
The paper ↗Extended patches through time as tubelets, and compared ways to factorise attention over space and time.
The paper ↗Researcher's shelf
Further reading
- Dosovitskiy et al., An Image is Worth 16x16 Words (ViT, 2020): https://arxiv.org/abs/2010.11929
- Liu et al., Visual Instruction Tuning (LLaVA, 2023): https://arxiv.org/abs/2304.08485
- Liu et al., Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, 2023): https://arxiv.org/abs/2310.03744
- Li et al., BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (2023): https://arxiv.org/abs/2301.12597
- Alayrac et al., Flamingo (2022): https://arxiv.org/abs/2204.14198
- Radford et al., Whisper (2022): https://arxiv.org/abs/2212.04356 and its code: https://github.com/openai/whisper
- Défossez et al., High Fidelity Neural Audio Compression (EnCodec, 2022): https://arxiv.org/abs/2210.13438
- van den Oord et al., Neural Discrete Representation Learning (VQ-VAE, 2017): https://arxiv.org/abs/1711.00937
- Chameleon Team, Chameleon: Mixed-Modal Early-Fusion Foundation Models (2024): https://arxiv.org/abs/2405.09818
- Arnab et al., ViViT: A Video Vision Transformer (2021): https://arxiv.org/abs/2103.15691
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.