rumblr Work in progressWIP

● The AI Primer · Lesson 37 · Generating images, audio and video

Multimodal models

images, audio and video into a language model

You'll be able to explain Images, audio and video into a language model

Members · open during launch 53 min15 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. One idea: every modality becomes a sequence of vectors as wide as the language model's word vectors; after that it is ordinary attention.
  2. Images: a Vision Transformer cuts a picture into patches, projects each one, adds a position, and runs non-causal encoder blocks. Tokens = (H/P)·(W/P), so they grow with the square of the resolution.
  3. Connecting to a language model: a small projector maps image vectors into the embedding space, and they are spliced into the prompt. Train the projector first with everything frozen, then fine-tune on image instructions.
  4. Audio: waveform, then a log-mel spectrogram (a Fourier transform per 25 ms slice, pooled into ear-shaped bands), then an encoder: about 50 tokens a second.
  5. Generating pictures and sound: snap vectors to a learned codebook so they become ids a language model can predict like words; or hand off to a diffusion model.
  6. Video and cost: frames multiply image tokens by time; sample frames, merge tokens, and remember that text is by far the densest input.

Level 1

The practitioner's guide

In one sentence

A multimodal model turns images, audio and video into sequences of vectors the same width as a language model's word vectors, so one transformer can read (and, with a codebook, write) all of them; for a practitioner, every picture, second of sound or frame of video is a number of tokens, and that number is the cost, the latency and the limit.

When you need it

You need this lesson the moment a model has to read something that isn't text: a screenshot, a chart, a scanned form, a voice message, a meeting recording, a video clip. Whether you call a hosted vision model or build with an open one, the same questions decide the outcome: how many tokens each input becomes, at what resolution, in what order in the prompt, and whether the model needs the signal at all or only its words. You don't need a multimodal model when the words are all that matters: a transcript of speech is about 3.2 tokens a second where audio tokens are 50 a second (this lesson's rates), and a caption is a few dozen tokens where a 336 × 336 image is 576. The number that shows how fast the naive approach fails: send a 10-second clip at 30 frames a second with 256 tokens a frame and it is 76,800 tokens, more than half of a 128,000-token window, before a single word of the question. Keep one frame a second and it is 2,560.

Your options

From the cheapest to the most control:

Option What it does What it guarantees What it costs Where it lives
Turn the signal into text first Transcribe audio (Whisper), caption or OCR an image, then use a text model The densest input there is: hours of speech in one window Loses tone, layout, small detail and anything the captioner missed Your pipeline, ahead of the model
A hosted multimodal API Send the image or audio in the prompt; the vendor's encoder tokenizes it No infrastructure; a documented token count per image Tokens per image by area (a 1000 × 1000 image is 1,296 tokens on one API); resizing above a size limit The vendor's API
An open vision-language model A ViT's patch vectors pass through a projector into the prompt (LLaVA's recipe) Full control of resolution, tiling and prompt layout A GPU; 196 to 576 or more tokens per image at your chosen resolution Your server
A compressed-visual family New cross-attention layers (Flamingo) or a small set of learned queries (BLIP-2's Q-Former) read the image instead of splicing it A short text sequence however many images; 32 to 64 tokens per image Some detail lost; a different model family to adopt The model family you download
Align your own projector Freeze a vision encoder and a language model, train the small adapter between them on captions, then tune on instructions A model that speaks your domain's pictures, cheaply (stage 1 runs in hours) Captioned data, then instruction data; a stage-2 fine-tune for behaviour Your training loop
Discrete tokens for generation Snap image or audio vectors to a learned codebook so the model predicts them like words One model reads and writes every modality with next-token prediction 1,024 ids for a 256 × 256 image; 600 ids a second for codec audio A tokenizer plus the language model

How to choose

Start from what the model must get out of the signal, then count the tokens.

  • Only the words matter (a voicemail, a lecture, a document with plain text): transcribe or OCR to text and use a text model. It is the densest and the cheapest by a wide margin.
  • Layout, tone, colour or small detail matters (a chart, a screenshot, a form, a hesitant customer): send the signal itself, at the smallest resolution that still shows the detail, cropped to the region that matters.
  • Long video: sample frames (one a second for a lecture, more for anything fast), merge neighbouring patch tokens, and send the soundtrack as a transcript alongside.
  • Your own domain, weak results from general models: align a projector on your captions first (cheap and safe, since both big models stay frozen), then instruction-tune.
  • Generating pictures or speech from a language model: discrete codebook tokens for simplicity, or hand the language model's output to a diffusion model (primer.ml.generative.diffusion) for finer detail.
  • Whatever you pick, put the image before the question. The language model is causal, so text placed before an image is scored as if the image were not there; the hosted API's own guidance says the same.

What it costs

Tokens grow with the square of the side: a 224 × 224 image in 16-pixel patches is 196 tokens, a 336 × 336 image in 14-pixel patches is 576, and a 1008 × 1008 image with 2 × 2 neighbours merged is 1,296. Hosted APIs bill the same way: on Claude's API each 28 × 28 pixel block is a visual token, so a 1000 × 1000 image is 1,296 tokens, larger images are downscaled to a long-edge limit (1568 pixels on the standard tier), and at $1 per million input tokens that image costs about $1.30 per thousand images. Audio is 50 encoder tokens a second on Whisper's design (a log-mel spectrogram at 100 frames a second, halved by a stride-2 convolution), so a 128,000-token window holds about 43 minutes; speech as text holds about 11 hours; video sampled at one frame a second and 256 tokens a frame holds about 8 minutes, and full-rate video under 20 seconds. Every one of those tokens lengthens the prefill before the first word of the answer and competes with instructions and retrieved documents for attention. Training cost is lopsided: stage-1 alignment trains a projector of a few million numbers between models of billions (this lesson's toy takes 80 pictures from 24% to 99% correct), and BLIP-2 beat the 80-billion-parameter Flamingo on zero-shot VQAv2 by 8.7% with 54 times fewer trainable parameters; it is stage 2, the instruction data, that decides how the model behaves.

What breaks

  • Objects that aren't there. A weakly aligned model describes what pictures like this usually contain; the hosted API warns of the same with low-quality, rotated or very small images (under 200 pixels). Send clear images at a usable size, and verify anything that matters.
  • The image after the question. The question's tokens cannot attend forward to an image that follows them. Image first, then the question.
  • Unreadable small text. A screenshot downscaled to the size limit loses its small print. Crop to the region, or tile the page, rather than sending one huge image.
  • The context window full of frames. 20 minutes of video at one frame a second and 256 tokens a frame is 307,200 tokens, over twice a 128k window. Sample less often, merge tokens, transcribe the audio, or split the clip.
  • Sampling that misses the moment. One frame a second is fine for a lecture and useless for a golf swing. Match the frame rate to the speed of what you are looking for.
  • Transcribing away the signal. A transcript drops hesitation, tone and who spoke; a caption drops layout. Use text only when the words are enough.
  • A codebook that fits nothing. A random codebook is useless at every size (error 0.6 to 2.1 against 0.09 to 0.32 for a learned one in this lesson's sweep); codebooks are learned on the data, and a tokenizer from one domain misbehaves on another.

In the wild

The Vision Transformer (Dosovitskiy et al.) is the front end: a pure transformer over 16 × 16 patches that matched convolutional networks given enough data, and CLIP's ViT (Radford et al.) is the one most vision-language models start from. LLaVA (Liu et al.) is the encode, project and splice recipe, trained on instruction data generated with a language model, scoring 85.1% of GPT-4 on its own multimodal benchmark; Flamingo (Alayrac et al.) reads images through new cross-attention layers between frozen models and learns tasks from a few examples in the prompt; BLIP-2 (Li et al.) compresses each image through a Querying Transformer. Whisper (Radford et al.) is the audio side: an encoder over log-mel spectrograms and a text decoder trained on 680,000 hours of audio, used zero-shot for transcription. For generation, VQ-VAE and DALL-E turned images into codebook ids for a transformer, SoundStream and EnCodec do it for audio with stacked residual codebooks, and Chameleon trains one model on interleaved text and image tokens from the start. Hosted vision APIs expose the same arithmetic as a price list: Claude's vision documentation, the source of the numbers above, counts each 28 × 28 pixel block as a visual token and caps images by long edge and token count. Every paper is linked at the end of the lesson.

Go deeper

Level 2 builds each front end by hand: a 4 × 4 picture cut into four tokens, a Vision Transformer with its projector spliced into a tiny language model and aligned on 80 pictures, a Fourier transform and a log-mel spectrogram from eight samples upward, a codebook learned by k-means with residual stages, and the token arithmetic for video and the context budget. If you only needed to count the tokens and choose, you are done.

Level 2

How it works, from scratch

Imagine a brilliant reader who can take in only one thing: a long row of index cards, each card holding a short list of numbers. That is a language model. Every word it reads arrives as one card, the word's vector (see primer.ml.tokenization and primer.ml.transformer). It has never seen a photograph or heard a voice.

To tell this reader about a photo, you cut the photo into small squares and write one card per square. To tell it about a voice recording, you write one card for every fiftieth of a second of sound. If the cards are written in the same "handwriting" as the word cards, the reader handles a photo the way it handles a sentence: it pays attention across all the cards at once.

That is the whole idea of a multimodal model, a model that takes in more than one modality (kind of input: text, images, audio, video):

Turn every modality into a sequence of vectors ("tokens") that a transformer can attend over.

Everything in this lesson is a way of making those cards: for images, for sound, for video, and, run backwards, for a model that produces pictures and speech.

Figure 1 · Diagram

Reading it: three roads, one destination. Each modality has its own front end (top three rows), and every front end ends in the same place: a list of vectors as wide as the language model's word vectors. From the box "One sequence of vectors" onwards there is no difference between a word, a patch of an image and a slice of sound; the language model attends over all of them together. The last box shows the trick for output: if pictures and sounds can be written as numbered tokens, the model can generate them the way it generates words (see "Generating images and speech" below).

Chapter 1

A tiny worked example: a 4×4 picture becomes 4 tokens

Take a grayscale picture of 4×4 pixels, each pixel a brightness from 0 to 15:

 0  1 |  2  3
 4  5 |  6  7
------+------
 8  9 | 10 11
12 13 | 14 15
  1. Cut it into 2×2 patches (the lines above): 4 patches.
  2. Flatten each patch into a list, reading row by row: the top-right patch becomes (2, 3, 6, 7).
  3. Project each list to the model's width, here 2 numbers, by multiplying by a matrix whose first column adds up the patch's top row and whose second column adds up its bottom row.
  4. Add a position: the patch's (row, column) in the grid, so the model knows where each patch came from.
Patch Pixels Projected + position Token
top left 0, 1, 4, 5 (1, 9) (0, 0) (1, 9)
top right 2, 3, 6, 7 (5, 13) (0, 1) (5, 14)
bottom left 8, 9, 12, 13 (17, 25) (1, 0) (18, 25)
bottom right 10, 11, 14, 15 (21, 29) (1, 1) (22, 30)

The picture is now four tokens of two numbers each: exactly the shape of input a transformer reads. A real model does the same with bigger numbers: patches of 14 or 16 pixels, and hundreds or thousands of numbers per token.

In code: worked_example_patches runs these four steps on this picture and returns the patches, the projections and the tokens.

Chapter 2

Images: a Vision Transformer from scratch

primer.ml.cnn_rnn introduced the idea of patches as tokens. Here we build the whole front end, the Vision Transformer (ViT), and count what it costs.

Cutting and counting

Everyday picture Lay a sheet of graph paper over a photo and cut along every fourth line. You get a grid of small tiles, and you can hand them over one by one, left to right and top to bottom, like the words of a sentence.

Tiny example A 224×224 image in 16-pixel patches is a grid of 224 / 16 = 14 patches down and 14 across: 14 × 14 = 196 tokens. A 336×336 image in 14-pixel patches is 24 × 24 = 576 tokens.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the image's height and width, in pixels 224, 224
the side of one square patch, in pixels 16
how many patches fit down the image (must be a whole number, so images are resized first) 14
how many patches fit across 14
multiply
the number of tokens the image becomes 196

Each token starts life as numbers, where is the number of colour channels (3 for red, green, blue): 16 · 16 · 3 = 768.

In words: "the number of image tokens is the number of patches down times the number of patches across."

With the numbers: (224 / 16) · (224 / 16) = 14 · 14 = 196; (336 / 14) · (336 / 14) = 24 · 24 = 576; the worked 4×4 picture in 2-pixel patches gives 2 · 2 = 4.

In Python:

H, W, P = 224, 224, 16
# H/P patches down, times W/P across
(H // P) * (W // P)  # → 196
H, W, P = 336, 336, 14
(H // P) * (W // P)  # → 576
# the worked 4×4 picture in 2-pixel patches
(4 // 2) * (4 // 2)  # → 4

Figure 2 · Drawn from the lesson's code

32×32 image, 8-pixel patches 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 0 10 20 30 40 50 60 the patch's 64 pixels, read row by row 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 token (patch number) …becomes 16 tokens of 64 numbers

A 32 by 32 picture of a sun over a striped field, cut into 16 numbered 8-pixel patches, and the 16-row matrix those patches become

Reading it: on the left, the orange lines cut a 32×32 picture into a 4 × 4 grid, numbered in reading order. On the right, each numbered patch has become one row: its 64 pixels laid end to end. Rows 0 to 7 (the sky, with the sun in patches 2, 3, 6 and 7) are mostly one grey; rows 8 to 15 (the striped field) repeat the same light and dark pattern. Nothing about the picture is lost, but it is now a list of 16 tokens, the same shape as a 16-word sentence.

Figure 3 · Drawn from the lesson's code

200 400 600 800 1000 1200 1400 image side length (pixels) 0 2000 4000 6000 8000 tokens per image Double the side, quadruple the tokens 336 px → 576 tokens 14-pixel patches 16-pixel patches 14-pixel patches, 2×2 merged

Tokens per image rise with the square of the side length: 576 at 336 pixels, over 9,000 at 1,344 pixels with 14-pixel patches, and a quarter of that when 2 by 2 neighbours are merged

Reading it: the x-axis is the side length of a square image, the y-axis the tokens it becomes. Each curve bends upward because the count grows with the square of the side: double the side and the tokens quadruple. Bigger patches (blue) give fewer tokens than smaller ones (red). The green curve merges each 2×2 block of neighbouring patch vectors into one token before the language model sees them, a common trick that cuts the count by four.

In code: image_to_patches cuts and flattens an image (the cutting is primer.ml.cnn_rnn.patchify), refusing sizes the patch does not divide, and count_image_tokens is the formula above, with an optional merge.

Why it matters in practice. Resolution is a cost dial. Small text in a screenshot needs a high resolution to be legible, and every doubling of the side costs four times the tokens (and attention cost grows faster still; see primer.ml.attention). Systems resize images to a supported size, cut very large ones into tiles, and merge neighbouring patches to keep the count manageable.

From patch to token: projection and position

Everyday picture Every tile gets the same questionnaire: "how bright is your top half? your bottom half? is there an edge?" The answers become the tile's card. Then a sticker with the tile's grid address goes on the card, so shuffling the cards loses nothing.

Tiny example In the worked example, the questionnaire had two questions (top-row total, bottom-row total), and the sticker was the patch's (row, column).

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape In the example (top-right patch)
which patch, counting in reading order from 0 1
patch flattened into a row of pixels (2, 3, 6, 7)
the learned patch-embedding matrix: one column per output number (a bias vector is usually added too; it is left out here) the 4 × 2 matrix below
a matrix multiply: each output number is the dot product of the patch with one column of (5, 13)
the learned position vector for slot (0, 1)
the finished token for patch (5, 14)
the model width: how many numbers per token 2

in the example is the rows (1, 0), (1, 0), (0, 1), (0, 1): the first two pixels (the patch's top row) feed output 1, the last two (its bottom row) feed output 2.

In words: "each token is its patch multiplied by a learned matrix, plus a learned vector that marks where the patch sits."

With the numbers: (2, 3, 6, 7) · = (2 + 3, 6 + 7) = (5, 13); adding the position (0, 1) gives (5, 14).

In Python:

x_i = [2, 3, 6, 7]
W_E = [[1, 0], [1, 0], [0, 1], [0, 1]]
p_i = [0, 1]
# x_i W_E: dot the patch with each column of W_E
xW = [sum(x * W_E[r][c] for r, x in enumerate(x_i)) for c in range(2)]
xW  # → [5, 13]
# + p_i
[a + b for a, b in zip(xW, p_i)]  # → [5, 14]

Why the position vector? Attention compares every token with every other and ignores their order (see primer.ml.positional). Without , a sky patch at the top and the same sky colour in a puddle at the bottom would be the same token, and "the sun is above the field" could not be expressed.

Figure 4 · Diagram

Reading it: follow the shapes under each box. Only the first two boxes are new; from "Encoder blocks" on, this is the transformer block of primer.ml.transformer, run with the causal mask switched off. A sentence is read left to right, so a text decoder hides the future; a picture has no future, so every patch may look at every other, above, below and to either side. The output has one vector per patch, now informed by the whole image.

In code: VisionEncoder holds W_E, the position table and a stack of primer.ml.transformer.TransformerBlock with causal=False; VisionEncoder.embed is the formula above, and calling the encoder runs the blocks.

Why it matters in practice. The vision encoder inside most vision-language models is a ViT that was first trained as the image half of CLIP (see primer.ml.embeddings.contrastive). CLIP training pulled its image vectors towards the vectors of matching captions, so its outputs already carry meaning a language model can use. That is why builders start from it instead of training vision from scratch.

Chapter 3

Connecting vision to a language model

The projector: a plug adapter

Everyday picture Your laptop charger has the right voltage but the wrong plug for the wall socket abroad. A small adapter changes the shape, not the power. The vision encoder speaks in vectors of its own width and style; the language model expects vectors shaped like its word embeddings. A projector is the adapter between them.

Tiny example A vision vector with 2 numbers, (1, 2), must become a language-model vector with 3 numbers.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here Shape In the example
one image token from the vision encoder (1, 2)
the projector's learned matrix rows (1, 0, 1) and (0, 1, 1)
the same token, now shaped like a word embedding (1, 2, 3)
the widths of the two models 2 and 3

In words: "multiply each image vector by one learned matrix to turn it into a vector the language model can read."

With the numbers: (1, 2) · = (1·1 + 2·0, 1·0 + 2·1, 1·1 + 2·1) = (1, 2, 3).

Level 3: in Python
v = [1, 2]
W_P = [[1, 0, 1], [0, 1, 1]]
# v W_P: dot v with each column of W_P
[sum(v[r] * W_P[r][c] for r in range(2)) for c in range(3)]  # → [1, 2, 3]

Many models use two such layers with a GELU between them (a smooth "keep the positives" function; see primer.ml.transformer), which is a small feed-forward network. Either way, the projector is tiny next to the two models it joins: a few million numbers between models with billions.

In code: Projector is a single linear layer by default and a two-layer MLP when given a hidden width; worked_example_projection computes the (1, 2, 3) above.

Splicing image tokens into the prompt

Everyday picture Writing a letter and taping a strip of photos into the middle of a sentence. The reader reads the words, then the photos, then the rest of the words, in order.

Tiny example The prompt "what is this <image>" has four text tokens, one of them a placeholder. The placeholder is replaced by the image's 4 projected vectors, so the language model reads 3 + 4 = 7 rows.

Step Shape (toy model) Meaning
image (8, 8) an 8×8 grayscale picture
patches (4, 16) four 4×4 patches, flattened
vision encoder output (4, 16) four image vectors of width 16
projector output (4, 24) four image tokens, as wide as the language model
text token vectors (3, 24) "what", "is", "this" from the embedding table
spliced sequence (7, 24) text, then image, in prompt order
logits (7, 12) a score for each of the 12 vocabulary words, at every position

Figure 6 · Diagram

Reading it: two front ends meet in the "Splice" box, where the placeholder's single row is replaced by the four projected image rows. After that the language model runs completely unchanged: positions are added (image tokens take up positions too), the causal decoder blocks run, and every row produces scores over the vocabulary. The language model never learns that some rows came from a picture.

Because the decoder is causal, a token can use the image only if it comes after the image. Text before the picture is scored exactly as if there were no picture. That is why prompts usually put the image first and the question after it.

In code: VisionLanguageModel.splice builds the spliced sequence and a flag for each row saying whether it is an image token; calling VisionLanguageModel runs primer.ml.transformer.TinyGPT's blocks on it; build_toy_vlm wires up the toy model in the table.

Why it matters in practice. This "encode, project, splice" design (as in LLaVA) is the simplest and most common. Two other families exist. One adds new cross-attention layers inside the language model that look at the image vectors (as in Flamingo), leaving the text sequence short. Another first compresses any image into a fixed small number of tokens, such as 32 or 64, with a little attention module of learned queries (as in BLIP-2's Q-Former). Both trade some detail for fewer tokens.

How such a model is trained

Everyday picture Two experts who don't share a language: a photographer and a writer. You hire an interpreter. First, the interpreter learns vocabulary while both experts carry on exactly as they are: the photographer points at pictures, the writer names them. Then all three practise real conversations together, and the writer is allowed to adapt a little too.

That is the usual two-stage recipe:

  1. Alignment. Freeze the vision encoder and the language model. Train only the projector on image-caption pairs, so the projected image tokens make the language model produce the caption.
  2. Visual instruction tuning. Train the projector and the language model (fully, or with LoRA; see primer.ml.training_stages) on images paired with instructions and answers: "what is unusual about this picture?" followed by a good answer.

In both stages the loss is ordinary next-token cross-entropy (the negative log of the probability given to the right token; see primer.ml.losses), counted only on the answer tokens. The model is not graded on predicting the image tokens or the question it was given.

Tiny example The answer is "horizontal stripes", two tokens. The model gives the right first token probability 0.5 and the right second token 0.25.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the positions of the answer tokens (the loss mask keeps only these) 2 positions
how many answer tokens there are 2
"for each position in the answer"
the correct token at position "horizontal", then "stripes"
every token before position
the probability the model gives the correct token, given () the image and what came before 0.5, 0.25
the natural logarithm; is 0 when and grows as shrinks (see primer.notation)
the loss: the average surprise over the answer 1.040

In words: "the loss is the average, over the answer tokens only, of how surprised the model was by each correct token."

With the numbers: −(log 0.5 + log 0.25) / 2 = (0.693 + 1.386) / 2 = 1.040.

Level 3: in Python
import math
# probabilities of the right answer tokens
p = [0.5, 0.25]
# −(1/|A|) Σ log p
round(-sum(math.log(q) for q in p) / len(p), 3)  # → 1.04

Figure 7 · Diagram

Reading it: the same three boxes appear in both stages; what changes is which ones learn. In stage 1 only the small middle box moves, so it is cheap and it cannot damage what the two big models already know. In stage 2 the language model joins in, so it learns to use the image tokens to follow instructions, answer questions and describe details, not just name things. The vision encoder often stays frozen throughout.

Our toy does stage 1. Its vision encoder and language model are random and frozen; only the projector learns, from 80 small pictures of four patterns (horizontal, vertical, diagonal, checkered), to make the language model's output layer pick the right pattern word.

Figure 5 · Drawn from the lesson's code

0 50 100 150 200 250 300 training step 0.5 1.0 1.5 2.0 2.5 caption cross-entropy Only the projector learns blind guess over 12 words 0 50 100 150 200 250 300 training step 0.0 0.2 0.4 0.6 0.8 1.0 right caption word ranked first Caption accuracy: 24% → 99% on new images training images chance (4 classes) new images

Stage-1 alignment: the caption loss falls from 2.41 to 0.33, and on 80 new images the right word ranks first 99% of the time, up from 24%

Reading it: on the left, the loss starts near the dashed line, the loss of a blind guess among 12 words (log 12 = 2.48), and falls steadily. On the right, the green line is the share of training images whose correct word scores highest, and the red dots are the same measure on 80 images the projector never saw: 24% before training (chance is 25%) and 99% after. Nothing but the projector changed, which is the point of stage 1: a small adapter is enough to make an existing vision encoder and an existing language model understand each other.

In code: align_projector trains only the projector's weights with cross-entropy on the caption word and returns the loss and accuracy per step; caption_accuracy measures how often the right word wins. The toy scores a pooled image vector directly against the language model's token table (the last step of the real path) rather than back-propagating through every layer.

Why it matters in practice. Stage 1 is cheap enough to run in hours, because almost every weight is frozen. Stage 2 decides the model's behaviour: the quality and variety of the image-instruction data matter more than its size. A common failure of a weakly aligned model is describing objects that are not in the picture, a visual form of hallucination.

Chapter 4

Audio: from a waveform to tokens

Sound is a list of numbers

Everyday picture A microphone is a tiny eardrum. It measures air pressure many thousands of times a second, and the recording is just that list of measurements: the waveform. Drawn on paper, it looks like a seismograph trace.

Tiny example Speech is usually recorded at a sample rate of 16,000 measurements per second, so one second is 16,000 numbers and a 30-second clip is 480,000. Used directly as tokens that would be ruinous, and the thing that matters, which pitches are sounding, is hidden in the wiggles (top panel of the spectrogram figure below).

How much of each pitch: the Fourier transform

Everyday picture A prism splits white light into its colours. The Fourier transform splits a sound into its pitches. It works by holding a pure wave of each pitch up against the recording and asking "how well do you line up?". A pitch that is present lines up again and again and scores high; a pitch that is absent lines up as often as it clashes and scores zero.

Tiny example Eight samples of a wave that goes up and down twice: (1, 0, −1, 0, 1, 0, −1, 0). Line it up against a wave that also cycles twice, cos: (1, 0, −1, 0, 1, 0, −1, 0). Multiply matching positions and add: 1 + 0 + 1 + 0 + 1 + 0 + 1 + 0 = 4. A wave that cycles once agrees for half its length and disagrees for the other half: it scores 0.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the -th sample of the recording (1, 0, −1, 0, 1, 0, −1, 0)
how many samples 8
a counter over the samples, from 0 to 0, 1, …, 7
which pitch we are testing: a wave that completes cycles in the samples 2
the cosine and sine waves (sine is the same wave shifted a quarter cycle, so a pitch that starts at a different moment is still caught)
one full cycle, in the units cos and sin use
add up over every sample
the length of the pair (cosine score, sine score)
how much of pitch the recording holds 4

Bin is the frequency in hertz (cycles per second), where is the sample rate. If the 8 samples were taken over one second, and bin 2 is 2 Hz. Textbooks write the same thing with complex numbers, ; the cosine part and the sine part are exactly the two sums above.

In words: "for each pitch, correlate the recording with a cosine and a sine of that pitch, and take the length of the two scores."

With the numbers: for the cosine sum is 4 and the sine sum is 0, so . For both sums are 0. Bin 6 also scores 4: with only 8 samples, a wave cycling 6 times is indistinguishable from one cycling 8 − 6 = 2 times, so the top half of the bins mirrors the bottom half and only bins 0 to are kept.

In Python:

import math
x = [1, 0, -1, 0, 1, 0, -1, 0]
N = len(x)
def magnitude(k):
    c = sum(x_n * math.cos(2 * math.pi * k * n / N) for n, x_n in enumerate(x))
    s = sum(x_n * math.sin(2 * math.pi * k * n / N) for n, x_n in enumerate(x))
    return math.sqrt(c ** 2 + s ** 2)
[round(magnitude(k), 6) for k in range(N)]  # → [0.0, 0.0, 4.0, 0.0, 0.0, 0.0, 4.0, 0.0]
# f_k = k · f_s / N, with 8 samples a second
2 * 8 / N  # → 2.0

In code: dft_magnitudes is this formula for every k at once, written from the definition (the tests check it against NumPy's fast Fourier transform, which computes the same numbers far faster).

The spectrogram: one Fourier transform per slice

Everyday picture A piano roll, or sheet music: time runs left to right, pitch runs bottom to top, and a mark means "this note is sounding now". To make one from a recording, listen through a short window, say 25 milliseconds, ask "which pitches are in here?", slide the window along a little, and ask again. That is the short-time Fourier transform (STFT), and its picture is a spectrogram.

Each slice is faded in and out with a Hann window (a smooth hump from 0 up to 1 and back) before measuring, because chopping a wave off abruptly creates a click, and a click contains every pitch at once.

Tiny example Speech systems commonly use 25 ms windows (400 samples at 16 kHz) that start every 10 ms (160 samples). How many windows fit in one second?

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the recording's length, in samples 16,000
the window length, in samples 400 (25 ms)
the hop: how far the window moves each step 160 (10 ms)
floor: round down to a whole number (a part-window at the end is dropped)
the number of frames (columns of the spectrogram) 98

In words: "one window fits at the start; after that, count how many whole hops still leave room for a full window."

With the numbers: 1 + ⌊(16,000 − 400) / 160⌋ = 1 + ⌊97.5⌋ = 98 frames per second, which is why speech models talk about "100 frames a second".

Level 3: in Python
L, N, H = 16000, 400, 160
# 1 + ⌊(L − N) / H⌋
1 + (L - N) // H  # → 98

Figure 10 · Diagram

Reading it: a flat list of samples becomes a grid. The first three boxes are the STFT; the shape under the Fourier box says each frame now holds one number per frequency bin. The last two boxes, explained next, shrink those bins into a few dozen bands spaced the way ears hear, and put loudness on a log scale. The result, a log-mel spectrogram, is what speech models actually read.

Figure 8 · Drawn from the lesson's code

0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 time (milliseconds): the first 20 ms only −1 0 1 pressure Waveform: 8,000 numbers a second, pitch hidden in the wiggles 0.0 0.2 0.4 0.6 0.8 1.0 time (seconds) 0 1000 2000 3000 4000 frequency (Hz) Spectrogram: 122 frames × 129 frequency bins 0.0 0.2 0.4 0.6 0.8 1.0 time (seconds) 0 10 20 30 40 mel band Log-mel spectrogram: 122 frames × 40 bands

A 20-millisecond waveform of wiggles, then a spectrogram of the same second of sound showing a flat line at 1000 Hz and a line rising from 200 to 3000 Hz, then the log-mel version where the rising line curves

Reading it: the signal is a whistle sliding from 200 Hz to 3000 Hz over a steady, quieter 1000 Hz hum, sampled 8,000 times a second. Top: the first 20 ms of the waveform, where both sounds are tangled into one wiggle. Middle: the spectrogram, time across and frequency up, brighter meaning louder. The two sounds separate cleanly: a flat line for the hum and a straight rising line for the whistle. Bottom: the log-mel version with 40 bands. The rising line now curves, climbing fast through the low bands and slowly through the high ones, because mel bands are narrow at low pitch and wide at high pitch, which is also where ears are more and less sensitive.

In code: stft slices, windows (with hann) and measures every frame; frame_count is the formula above; log_mel_spectrogram adds the mel pooling and the log.

The mel scale: spacing pitch the way ears do

Everyday picture On a piano, every octave takes up the same width of keyboard, yet each octave doubles the frequency: the A keys are at 110, 220, 440 and 880 Hz. Ears work the same way above about 1,000 Hz: a jump from 1,000 to 2,000 Hz sounds about as big as a jump from 2,000 to 4,000. The mel scale relabels frequencies so that equal steps in mel sound like equal steps in pitch.

Tiny example By construction 1,000 Hz is 1,000 mel. The 7,000 Hz from 1,000 to 8,000 Hz shrinks to just 1,840 mel.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
a frequency in hertz 700
frequency measured in units of 700 Hz; below about 700 Hz the scale is nearly straight, above it the log takes over 1
the base-10 logarithm: "10 to what power gives this?"; it turns ratios into equal steps (see primer.notation)
2595 a constant chosen so that 1,000 Hz comes out as 1,000 mel
the same frequency in mel 781.2

In words: "mel is a log of the frequency, gently straightened below 700 Hz and scaled so that 1,000 Hz is 1,000 mel."

With the numbers: 2595 · log₁₀(1 + 700/700) = 2595 · 0.301 = 781.2; 1,000 Hz gives 1,000.0; 4,000 Hz gives 2,146.1; 8,000 Hz gives 2,840.0.

Level 3: in Python
import math
def mel(f):
    return 2595 * math.log10(1 + f / 700)
round(mel(700), 1)  # → 781.2
round(mel(1000), 1)  # → 1000.0
round(mel(4000), 1)  # → 2146.1

A mel filterbank turns the hundreds of frequency bins into a few dozen bands: triangles evenly spaced in mel, so narrow at low frequencies and wide at high ones. Each band adds up the energy under its triangle. Loudness is also heard by ratio, so the last step takes the log.

Figure 9 · Drawn from the lesson's code

0 2000 4000 6000 8000 frequency (Hz) 0 2000 4000 6000 8000 mel Mel: even steps low, squeezed high 1000 Hz = 1000 mel if mel were hertz 0 2000 4000 6000 8000 frequency (Hz) 0.0 0.2 0.4 0.6 0.8 1.0 weight 10 mel filters: narrow low, wide high

Left: the mel curve rises steeply to 1000 mel at 1000 Hz and then flattens, reaching only 2840 at 8000 Hz. Right: ten triangular filters, narrow below 1000 Hz and ever wider above

Reading it: on the left, the purple curve is the formula and the dashed line is where it would be if mel were simply hertz. They agree near the bottom (the red dot, 1,000 Hz = 1,000 mel) and part ways above: the top 7,000 Hz of the range is squeezed into under 2,000 mel. On the right are 10 mel filters for 16 kHz audio. Each peaks at 1 and overlaps its neighbours; the low ones are a few bins wide and the top one spans over 3,000 Hz. Detail is spent where ears, and speech, need it.

In code: hz_to_mel is the formula, and mel_filterbank builds the triangles, evenly spaced in mel.

Speech recognition: an encoder over frames, a decoder for text

Everyday picture A court stenographer listens to a whole sentence, then types it out word by word, glancing back at what they heard as they go.

Tiny example Take a 30-second clip. With 10 ms hops it is 3,000 frames of 80 mel bands each. A small convolution that moves two frames at a time halves that to 1,500 encoder positions: 50 audio tokens per second. A transformer encoder attends over all 1,500 at once, and a decoder writes the transcript as text tokens, attending both to its own words so far and to the encoder's output. This is the design of Whisper.

Figure 11 · Diagram

Reading it: the left half is a front end like the image one: turn the signal into a grid, shrink it, and run a non-causal encoder so every moment of sound can inform every other. The right half is a text decoder with one extra step, cross-attention, where its queries come from the text written so far and its keys and values come from the audio (see primer.ml.transformer for encoders and decoders). The dotted arrow is the generation loop: each new word is fed back in.

In code: audio_tokens counts encoder positions from a clip's length, the hop and the downsampling.

Why it matters in practice. The same audio encoder can feed a general language model instead of a dedicated decoder: add a projector, splice the audio tokens into the prompt, and the model can answer questions about a recording, exactly as with images. Fifty tokens a second is manageable for a short voice message and expensive for an hour-long meeting.

Chapter 5

Generating images and speech: discrete tokens from a codebook

So far the model reads pictures and sound. To write them, a language model needs them as something it can predict one at a time from a fixed vocabulary, the way it predicts words.

Everyday picture A paint-by-numbers kit comes with a palette of numbered pots. Any small area of a picture is described by the number of the pot closest to its colour. The whole picture becomes a grid of pot numbers, and anyone with the same palette can repaint it. The palette is a codebook; snapping to the nearest entry is vector quantization.

Tiny example A codebook of three 2-number entries: , , . The vector has squared distances 0.85, 0.05 and 1.45 to them, so it becomes id 1. Decoding id 1 gives back (1, 0): close, not exact.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
one vector to encode: an image patch's vector or a slice of sound (0.9, 0.2)
codebook entry number
the codebook size: how many entries 3
" ranges over the entry numbers" 0, 1, 2
squared distance: subtract, square each difference, add
"the that gives the smallest value" (the position of the minimum, not the minimum itself) 1
the id that replaces 1

In words: "replace each vector by the number of its nearest codebook entry."

With the numbers: distances squared to are 0.81 + 0.04 = 0.85, 0.01 + 0.04 = 0.05 and 0.81 + 0.64 = 1.45; the smallest is 0.05, at .

Level 3: in Python
z = [0.9, 0.2]
codebook = [[0, 0], [1, 0], [0, 1]]
# ‖z − c_k‖² for each entry
d2 = [sum((a - b) ** 2 for a, b in zip(z, c)) for c in codebook]
[round(d, 2) for d in d2]  # → [0.85, 0.05, 1.45]
# arg min: the position of the smallest
d2.index(min(d2))  # → 1

Figure 13 · Diagram

Reading it: the top row is used during training: real pictures and sounds are encoded, snapped to the codebook, and become sequences of ids that sit in the training text next to their descriptions. The language model learns to continue a caption with ids the same way it learns to continue a sentence with words; its vocabulary is simply enlarged by entries. At generation time (bottom row) the model writes ids, the codebook turns them back into vectors, and a decoder paints pixels or synthesises a waveform. The encoder, codebook and decoder together are a VQ-VAE (see primer.ml.generative.autoencoders).

How is the codebook chosen? It is learned so that the entries sit where the data actually is, which amounts to k-means clustering (see primer.ml.embeddings.clustering): snap every vector to its nearest entry, move each entry to the average of the vectors that chose it, and repeat. For audio, one codebook is rarely precise enough, so neural audio codecs use residual vector quantization: a second codebook encodes what the first one missed, a third what the second missed, and so on. Each moment of sound becomes a small stack of ids.

Figure 12 · Drawn from the lesson's code

0 2 4 6 8 10 12 bits per patch (log₂ of codebook size, summed over stages) 1 0 − 1 1 0 0 mean squared error per pixel A learned codebook, and a second one for the leftovers clean pattern, noise dropped (0.3² = 0.09) random codebook learned codebook (k-means) residual: 1 to 4 stages of 8 entries

On new patches, a learned codebook's error falls from 0.32 to about 0.09 as it grows to 64 entries while random codebooks stay far above it; stacking residual codebooks of 8 entries keeps pushing the error down

Reading it: the x-axis is bits per patch: a codebook of entries costs log₂ bits per id, so 64 entries is 6 bits. The y-axis (log scale) is the reconstruction error on patches of new pictures, not the ones the codebook was learned from. Random codebooks (red) are useless at every size: their entries sit where no data lives. Learned codebooks (blue) fall fast and then flatten near the dotted line, which is the error you would get by reproducing each clean pattern perfectly and dropping only the pixel noise. The green squares stack 8-entry codebooks, each learned on the previous one's leftovers: after four stages (12 bits) they reach below the dotted line, because the extra codes start describing the noise of each particular patch too.

In code: quantize is the arg min above and dequantize the lookup; kmeans_codebook learns a codebook; learn_residual_codebooks and residual_quantize stack codebooks on the leftovers.

Why it matters in practice. Discrete tokens let one model read and write every modality with one mechanism, next-token prediction. The costs are real: a 256×256 image compressed eight-fold per side is 32 × 32 = 1,024 ids (the scheme of the original DALL·E), and a neural audio codec running at 75 frames a second with 8 codebooks writes 600 ids a second. The other road to generating images is diffusion (see primer.ml.generative.diffusion): the language model supplies text or vectors that condition a diffusion model, which paints the pixels in many small denoising steps. Discrete tokens are simpler to bolt onto a language model; diffusion usually renders finer detail.

Chapter 6

Video: pictures over time

Everyday picture A flipbook. Each page is a picture, and neighbouring pages are almost identical. To follow the story you don't need to look at every page; a glance at every tenth page shows what happens.

Tiny example A 10-second clip at 30 frames a second, with each frame cut into 256 tokens (224×224 in 14-pixel patches): 300 frames × 256 = 76,800 tokens, more than half of a 128,000-token context, for ten seconds. Keeping one frame a second gives 10 × 256 = 2,560.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the clip's duration, in seconds 10
frames kept per second (the sampling rate, often far below the video's own frame rate) 30, then 1
the tubelet depth: how many consecutive frames share one token, when a patch is cut through time as well as space 1
how many frame slots become tokens 300, then 10
tokens per frame, from the image formula 16 · 16 = 256
the clip's total tokens 76,800, then 2,560

In words: "count the frames you keep, divide by how many frames each token spans, and multiply by the tokens per frame."

With the numbers: (10 · 30 / 1) · 16 · 16 = 300 · 256 = 76,800; at one frame a second, (10 · 1 / 1) · 256 = 2,560; sampling 2 frames a second in tubelets 2 frames deep gives (10 · 2 / 2) · 256 = 2,560 again, while covering twice as many moments.

Level 3: in Python
H = W = 224
P = 14
tokens_per_frame = (H // P) * (W // P)
tokens_per_frame  # → 256
D, r, t = 10, 30, 1
D * r // t * tokens_per_frame  # → 76800
# one frame a second
D, r, t = 10, 1, 1
D * r // t * tokens_per_frame  # → 2560

Figure 14 · Diagram

Reading it: each box from left to right cuts the token count. Frame sampling throws away near-duplicate pages of the flipbook; tubelets and merging squeeze each kept moment into fewer tokens. A short text timestamp before each frame's tokens ("at 0:05") lets the model say when something happened. What is left is the image pipeline, repeated once per kept frame.

In code: video_tokens is the formula, and sample_frames picks the evenly spaced frames to keep.

Why it matters in practice. Sampling is a bet that nothing important happens between the frames you keep: one frame a second is fine for a lecture and useless for a golf swing. Real systems adapt: more frames for short or fast clips, fewer and smaller ones for long recordings, and a transcript of the soundtrack alongside.

Chapter 7

What it costs: tokens per second and the context budget

Everyday picture Packing a suitcase. Text is neatly folded shirts; images are shoes; video is an inflatable boat. The suitcase (the context window) is the same size whatever you pack.

Tiny example A 128,000-token window holds 128,000 / 256 = 500 seconds, about 8 minutes, of video sampled at one frame a second, before a single word of the question. The same window holds about 11 hours of speech written out as text.

Input Tokens per second Minutes in 128k tokens
speech written as text (about 150 words a minute, about 1.3 tokens a word) 3.2 656
speech as audio-encoder tokens 50 43
video, 1 frame a second, 256 tokens a frame 256 8.3
audio as codec tokens (75 frames × 8 codebooks) 600 3.6
video, 30 frames a second, 256 tokens a frame 7,680 0.3

Figure 15 · Drawn from the lesson's code

1 0 0 1 0 1 1 0 2 1 0 3 1 0 4 length of the recording (seconds; 60 = a minute, 3600 = an hour) 1 0 1 1 0 2 1 0 3 1 0 4 1 0 5 1 0 6 1 0 7 1 0 8 tokens How fast each kind of input fills a context window 128k-token window 1M-token window speech written as text (150 words a minute) speech as audio-encoder tokens video, 1 frame a second, 256 tokens a frame audio as codec tokens (75 frames × 8 codebooks) video, 30 frames a second, 256 tokens a frame

On log axes, five straight lines of tokens against recording length: speech-as-text crosses a 128k window only after hours, 30-frames-a-second video within seconds

Reading it: both axes are logarithmic, so each line is the same slope shifted up or down by its tokens per second. Read across at a dashed line (a context window) to see how long a recording of each kind fits. Text (grey) is so dense that hours of speech fit easily. Audio tokens (green) fill a 128k window in about 43 minutes, sampled video (blue) in about 8, and full-rate video (red) in under 20 seconds. The gap between the grey line and the others is the price of keeping the raw signal instead of words.

In code: seconds_that_fit divides a window by a rate, and TOKENS_PER_SECOND holds the rates in the table.

Why it matters in practice. Every image and second of audio is billed and computed like text: it takes room in the context window, lengthens the prefill before the first word of the answer (see primer.ml.inference), and competes with instructions and retrieved documents for attention (see primer.agents.context). The practical levers follow from the formulas: send the smallest resolution that still shows the detail you need, crop to the region that matters, sample fewer frames, and transcribe audio to text when the words matter and the tone does not.

Test yourself

8 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1How can a chatbot "see" a photo, explained without jargon?Think it through, then reveal

The photo is cut into a grid of small squares, and each square is described by a list of numbers in the same format the model uses for words. The model then reads the squares and the words of your question together, paying attention to whichever squares help answer it. It never sees the picture as a picture: it reads it as a few hundred extra "words".

Question 2How many tokens is a 448×448 image with 14-pixel patches? And if each 2×2 block of neighbours is merged?Think it through, then reveal

448 / 14 = 32 patches a side, so 32 × 32 = 1,024 tokens. Merging 2×2 blocks divides by 4: 256 tokens. Doubling the side from 224 would have quadrupled the count.

Question 3Why does a Vision Transformer need position vectors, and why is its attention not causal?Think it through, then reveal

Attention ignores order, so without positions the model sees an unordered bag of patches and cannot tell the sky from the ground. It is not causal because a picture has no "future": every patch may use every other, and hiding half the image would only throw information away.

Question 4What does the projector do, and why train it first with both big models frozen?Think it through, then reveal

It maps each image vector into the language model's embedding space, so image tokens look like something the language model can use. Training it alone first is cheap (it is tiny) and safe (the frozen models keep all they know). Only once the two sides understand each other is the language model tuned on image instructions.

Question 5In a prompt, why does it matter whether the image comes before or after the question?Think it through, then reveal

The language model is causal: a token can only attend to tokens before it. Question tokens placed after the image can look at it while being processed; question tokens placed before it cannot. The answer comes after both either way, but putting the image first lets the whole question be read in light of the picture.

Question 6Why does a speech model read a log-mel spectrogram rather than raw samples?Think it through, then reveal

Raw audio is 16,000 numbers a second with pitch hidden in the wiggles. The spectrogram makes pitch explicit (which frequencies sound at each moment), the mel bands spend detail where ears and speech need it, the log matches how loudness is heard, and 100 frames a second is far fewer positions to attend over.

Question 7How can a model that only predicts tokens produce an image or a voice?Think it through, then reveal

Encode pictures or sounds into vectors, snap each vector to its nearest entry in a learned codebook, and use the entry numbers as extra vocabulary. The model learns to predict those ids after a caption, and a decoder turns generated ids back into pixels or a waveform. The alternative is to have the model condition a diffusion model that paints the image.

Question 8A 20-minute video goes to a model with a 128,000-token window. What goes wrong, and what can be done?Think it through, then reveal

At one frame a second and 256 tokens a frame it is 1,200 × 256 = 307,200 tokens, more than twice the window, before any question. Options: sample far fewer frames (one every 5 seconds gives 61,440), merge neighbouring patch tokens, lower the resolution, transcribe the soundtrack to text, or split the video and summarise the parts.

Primary sources

The papers behind this lesson

Dosovitskiy et al., An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (2020)

Showed that a plain transformer over image patches matches convolutional networks when trained on enough data: the Vision Transformer.

Read the annotated companion →The paper ↗
Radford et al., Learning Transferable Visual Models From Natural Language Supervision (CLIP, 2021)

Trained an image encoder and a text encoder into one shared space; its image encoder is the starting point of many vision-language models.

Read the annotated companion →The paper ↗
Alayrac et al., Flamingo: a Visual Language Model for Few-Shot Learning (2022)

Connected a frozen vision encoder to a frozen language model through new cross-attention layers, handling images and video interleaved with text.

Read the annotated companion →The paper ↗
Liu et al., Visual Instruction Tuning (LLaVA, 2023)

The encode, project and splice recipe, trained in two stages: align the projector, then tune on image instructions.

Read the annotated companion →The paper ↗
Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (Whisper, 2022)

An encoder over log-mel spectrograms and a text decoder, trained on 680,000 hours of transcribed audio.

Read the annotated companion →The paper ↗
van den Oord, Vinyals & Kavukcuoglu, Neural Discrete Representation Learning (VQ-VAE, 2017)

Learned a codebook inside an autoencoder, turning images and audio into discrete tokens.

Read the annotated companion →The paper ↗
Zeghidour et al., SoundStream: An End-to-End Neural Audio Codec (2021)

Residual vector quantization for audio: a stack of codebooks, each encoding the previous one's error.

The paper ↗
Ramesh et al., Zero-Shot Text-to-Image Generation (DALL·E, 2021)

Generated images as sequences of discrete codebook tokens after the text, with one transformer.

The paper ↗
Arnab et al., ViViT: A Video Vision Transformer (2021)

Extended patches through time as tubelets, and compared ways to factorise attention over space and time.

The paper ↗

Researcher's shelf

Further reading

  • Dosovitskiy et al., An Image is Worth 16x16 Words (ViT, 2020): https://arxiv.org/abs/2010.11929
  • Liu et al., Visual Instruction Tuning (LLaVA, 2023): https://arxiv.org/abs/2304.08485
  • Liu et al., Improved Baselines with Visual Instruction Tuning (LLaVA-1.5, 2023): https://arxiv.org/abs/2310.03744
  • Li et al., BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models (2023): https://arxiv.org/abs/2301.12597
  • Alayrac et al., Flamingo (2022): https://arxiv.org/abs/2204.14198
  • Radford et al., Whisper (2022): https://arxiv.org/abs/2212.04356 and its code: https://github.com/openai/whisper
  • Défossez et al., High Fidelity Neural Audio Compression (EnCodec, 2022): https://arxiv.org/abs/2210.13438
  • van den Oord et al., Neural Discrete Representation Learning (VQ-VAE, 2017): https://arxiv.org/abs/1711.00937
  • Chameleon Team, Chameleon: Mixed-Modal Early-Fusion Foundation Models (2024): https://arxiv.org/abs/2405.09818
  • Arnab et al., ViViT: A Video Vision Transformer (2021): https://arxiv.org/abs/2103.15691

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.