rumblr Work in progressWIP

● The AI Primer · Lesson 36 · Generating images, audio and video

Diffusion and flow matching

turning noise into data one small step at a time

You'll be able to explain Turning noise into images one small step at a time

Members · open during launch 50 min14 figures and diagrams
Guide is what to use and when. How it works builds it from scratch. Math & code adds the formulas and the Python.

The lesson in one minute

What you'll be able to explain

  1. Forward process: mix data with a little Gaussian noise per step until only noise is left; the shortcut jumps to any step in one go.
  2. Training: show the network a noised example and the step; it guesses the noise; grade by squared error. Plain, stable regression, and secretly learning the score (the direction toward the data).
  3. Sampling: start from pure noise and repeatedly remove a little guessed noise (DDPM, with a fresh wobble each step), or take a few big deterministic jumps (DDIM).
  4. Flow matching: learn the velocity along straight lines from noise to data and follow it with Euler steps; straighter paths need fewer steps.
  5. Guidance: ask the same network with and without the label and push past the labelled guess; more on-prompt, less varied.
  6. At scale: denoise an autoencoder's small latent with a transformer that reads a text encoder's tokens; video and audio are the same idea with more tokens.

Level 1

The practitioner's guide

In one sentence

A diffusion model generates by learning one modest skill, "remove a little noise from this noisy example", and running it many times starting from pure noise; flow matching is the same idea along straight paths with fewer steps; together they are the engine inside today's image, video and audio generators, and the settings you meet (steps, guidance scale, resolution, seed) are the dials of that engine.

When you need it

You need this lesson the moment you generate or edit images, video or audio: whether you call a hosted model or run an open one, the choices you make (which model family, how many steps, what guidance scale, what size, which sampler) are the ones below, and the bill and the artefacts follow from them. You don't need diffusion for text (language models generate one token at a time), for a one-pass generator in a real-time loop (a GAN or a distilled few-step model, see primer.ml.generative.gans), or for a tiny domain with a handful of factors (a VAE will do, see primer.ml.generative.autoencoders). The number that shows the naive approach failing: ask this lesson's trained model for a sample in a single step and every sample lands on the data's average, the empty middle between the four blobs (mean distance from the centre 0.08); give it 20 well-placed jumps and its samples sit as close to the data as a fresh draw of the data itself (0.056 against 0.055).

Your options

From the least commitment to the most:

Option What it does What it guarantees What it costs Where it lives
A hosted image, video or audio model Prompt in, sample out; you set steps, guidance, size and seed No infrastructure; the vendor's sampler and safety checks A price per sample, and only the dials the API exposes The vendor's API
An open latent-diffusion model in a pipeline A pretrained denoiser, text encoder and autoencoder you run Full control of sampler, steps, guidance, seed and adapters A GPU with enough memory; defaults of 50 steps and guidance 7.5 in diffusers Your server
A faster sampler on the same model DDIM or a higher-order solver takes big deterministic jumps The same trained weights, 10 to 50 times faster than the original 1,000 steps Quality falls off below about 10 calls A setting in the pipeline
A flow-matching or rectified-flow model Trained to follow straight paths from noise to data Fewer steps for the same quality (near the floor by 5 calls here, against 10 for diffusion) A model trained that way, such as Stable Diffusion 3 The model family you download
A distilled few-step model A student trained to match the full model in 1 to 4 steps, often with an adversarial loss Real-time sampling Some variety and detail; a teacher and a distillation run The fast path beside the full model
Your own diffusion model This lesson's training loop on your own data A generator for a narrow domain nobody has published Data, a training run (half a second for the lesson's toy; hundreds of GPU-days for pixel-space image models), an autoencoder if the data is large Your training loop

How to choose

Start from the latency you can afford, then set the dials in this order.

  • Steps first. The number of network calls is the price of a sample. Start at the pipeline default (50 in diffusers), halve it while the output holds, and reach for a flow model or a distilled one when you need fewer than about 10.
  • Guidance scale next. 0 ignores the prompt (27% of samples on the asked-for blob here), 1 follows it plainly (100% on target, with the data's own spread), and above 1 exaggerates it: at 3 every sample is on target but bunched into a knot at the far edge of the blob, spread 0.14 against the real 0.27. Text-to-image models default to well above 1 (7.5 in diffusers), because unguided samples follow prompts loosely; raise it for obedience, lower it for variety.
  • Resolution: generate at the size the model was trained on (Stable Diffusion 1.x was trained on 512 × 512 images), and upscale afterwards. Every doubling of side length quadruples the latent and the attention cost more than that.
  • Determinism: DDIM and flow samplers add no fresh noise, so one seed gives one image, which is what makes a seed reproducible and an image editable by re-running with a changed prompt.
  • Whatever you pick, judge the output on variety as well as quality. Every dial that makes samples match the prompt better makes them more alike.

What it costs

Latency is steps times the cost of one denoiser call, and guidance doubles the calls per step (the network runs once with the prompt and once without). Working in an autoencoder's latent cuts each call 48-fold for a 512 × 512 image (786,432 numbers down to 16,384), which is what made high-resolution diffusion affordable on one GPU. Video multiplies it back: 16 frames of the same latent are 16 times the tokens and, because attention compares every token with every other, 256 times the attention work. The denoiser itself is now usually a transformer over latent patches, and the DiT paper found that more compute per sample (a deeper or wider model, or more tokens) gives lower FID, down to 2.27 on 256 × 256 ImageNet. Training a real model is the expensive part: pixel-space diffusion "often consumes hundreds of GPU days" in the latent-diffusion paper's own words, which is why the pretrained autoencoder and the latent exist. The training loss itself is plain squared-error regression with no adversary, which is why the runs are stable and why diffusion displaced GANs.

What breaks

  • Too few steps. Below roughly 10 calls a diffusion sampler's output drifts toward the average (the empty middle here; blur or mush in images). Use a flow or distilled model instead of starving a diffusion one.
  • Guidance too high. Oversaturated, samey, exaggerated images, the image version of the tight knot at weight 3; the diffusers documentation puts it as prompt adherence "at the expense of lower image quality". Pull it back toward the default.
  • Guidance too low. The prompt is followed loosely or not at all, as at weight 0. A guidance scale of 1 or below switches guidance off.
  • The wrong size. A model asked for a resolution or aspect ratio far from its training size composes badly (repeated subjects, stretched scenes). Generate at the native size and upscale.
  • A mismatched autoencoder or scale. The latent must be decoded by the autoencoder the denoiser was trained with, with the library's scaling factor applied both ways, or the output is junk; see primer.ml.generative.autoencoders.
  • Averaging without the wobble. DDPM adds a small fresh noise each step so samples commit to one possibility; a sampler that only ever steps to the network's average drifts to the safe, blurry middle. Deterministic samplers avoid this by jumping along the predicted noise direction, not to the mean.
  • Cost that scales with frames. A video request costs attention quadratically in its token count; a short clip at a modest size is many images' worth of compute, and the bill follows.

In the wild

Stable Diffusion is the open reference: an 860-million parameter U-Net denoising a 64 × 64 × 4 latent, conditioned through cross-attention on a frozen CLIP ViT-L/14 text encoder, trained on 512 × 512 images, and run through Hugging Face diffusers with 50 steps and guidance 7.5 by default. DDIM (Song, Meng and Ermon) gave the 10 to 50 times faster deterministic sampler every pipeline offers; classifier-free guidance (Ho and Salimans) is the guidance-scale slider, trading variety for fidelity with no separate classifier; DiT (Peebles and Xie) replaced the U-Net with a transformer over latent patches and showed it scales; Stable Diffusion 3 (Esser et al.) trains a transformer as a rectified flow with 16-channel latents; Adversarial Diffusion Distillation turns a foundation model into a one-to-four-step sampler. Video generators run the same denoiser over patches that span space and time, and audio generators denoise a spectrogram or an audio autoencoder's latent. Every paper is linked at the end of the lesson.

Go deeper

Level 2 builds the whole engine on four blobs of dots: the forward process and its one-jump shortcut, the noise-guessing loss and why it is secretly learning the direction toward the data, DDPM and DDIM sampling step by step, flow matching along straight lines, guidance with its knot, and the latent-and-patches arithmetic behind real systems. If you only needed to set the dials, you are done.

Level 2

How it works, from scratch

Imagine a sculptor who cannot carve a statue in one go. Hand them a shapeless block and ask for a horse, and they freeze. But hand them a slightly rough horse, and they can always make it a little less rough: smooth this bump, deepen that line. Now chain that one modest skill. Start from a shapeless block, make it a bit more horse-like, then a bit more, a hundred times over, and a horse comes out.

A diffusion model is that sculptor. It never learns "draw a picture from nothing", which is very hard. It learns "here is a picture with a little too much noise in it: take some of the noise away", which is much easier. And the practice material is free: take any real picture, add noise yourself, and you know exactly what the answer should be.

Figure 1 · Diagram

Reading it: the top arrow is the forward process: it destroys data by adding noise a little at a time, and it involves no learning at all. The bottom arrow runs the same road backwards, and that direction is what the network learns. At generation time only the bottom arrow runs, starting from fresh noise, so every run ends at a different, new sample.

The tiny world this lesson works in. A real image is a long list of numbers, one per pixel colour. To keep everything small enough to see, our "images" have just two numbers, so each one is a dot on a flat plane. The real data is four round blobs of dots, one in each corner (north-east, north-west, south-west, south-east), 600 dots in all. The job: learn to make new dots that land in the blobs, starting from random noise. Every idea below works unchanged on a 512 × 512 colour image; it just has 786,432 numbers per "dot" instead of 2.

Notice one thing about the blobs now, because it matters later: their average is the empty middle. A model that plays safe and outputs "the average answer" puts its dots where no real data lives. The same failure makes blurry images: the average of many sharp faces is a blur.

Chapter 1

Step 1: the forward process, adding noise a little at a time

Everyday picture Think of an old television slowly losing its signal. Turn the static dial up one notch: the picture fades a little and a little snow creeps in. Turn it a hundred notches and you see only snow, and nothing in it tells you what the programme was.

Tiny worked example Follow one pixel whose clean value is . Each step keeps most of the signal and mixes in a little fresh random noise. The noise is a random number drawn from a standard normal distribution: the bell curve centred on 0 whose typical size is 1 (values like 0.5, −1.2 and 0.1 are common, 3 is rare). Say step 1 mixes in a share of noise, and the random draw is .

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the step number, from 1 to (the last step) 1
the value before this step ( is the clean data) 2.0
the value after this step 2.06
"beta": how much noise this step mixes in, a number between 0 and 1 0.1
how much of the old value survives: it shrinks a little
"epsilon": fresh noise drawn from the standard normal bell curve 0.5
how much noise is mixed in

In words: "shrink the old value a little, then add a little random noise."

With the numbers: $x_1 = \sqrt{0.9} \cdot 2.0 + \sqrt{0.1} \cdot 0.5 = 1.897 + 0.158 = 2.06$.

Level 3: in Python
import math
x_prev, beta, eps = 2.0, 0.1, 0.5
# √(1 − β) x_{t−1}: shrink the old value
round(math.sqrt(1 - beta) * x_prev, 3)  # → 1.897
# √β ε: the noise mixed in
round(math.sqrt(beta) * eps, 3)  # → 0.158
round(math.sqrt(1 - beta) * x_prev + math.sqrt(beta) * eps, 2)  # → 2.06

Why the square roots? A value's variance (its typical squared size) is what matters, and variances of independent things add. If the old value has variance 1, the new one has variance : exactly 1 again. The square roots keep every step at the same overall size, so after many steps the value is noise of size 1, not a number that has blown up or faded to zero.

The shortcut: jump straight to any step

Running steps one by one is slow, and training needs noisy examples at every step. Luckily, adding noise twice is the same as adding it once in a bigger dose, so there is a formula that jumps straight to step . First, multiply up how much signal survives every step so far:

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the share of variance step keeps (often written ) 0.9, then 0.8
"multiply together, for = 1, 2, …, " (see primer.notation)
"alpha bar": the share of the original signal's variance left after steps; 1 means clean, 0 means pure noise 0.72; or 0.64 in the jump below
the clean data 2.0
one fresh standard-normal draw, standing in for all the noise added so far 0.5
the signal share: how much of is still in
the noise share

In words: "multiply together the survival share of every step so far; then the noisy value is that much of the clean value plus the rest in noise."

With the numbers: two steps with then leave . In this lesson's schedule, step 30 leaves , so the pixel at 2.0 with noise 0.5 lands at .

Level 3: in Python
import math
betas = [0.1, 0.2]
alpha_bar = 1.0
# ∏ (1 − β_s): multiply up what each step keeps
for beta in betas:
    alpha_bar *= 1 - beta
round(alpha_bar, 2)  # → 0.72
# the jump to step 30, where ᾱ = 0.64
x0, eps, alpha_bar_30 = 2.0, 0.5, 0.64
round(math.sqrt(alpha_bar_30) * x0 + math.sqrt(1 - alpha_bar_30) * eps, 2)  # → 1.9

Figure 4 · Diagram

Reading it: the solid chain is the slow way: one small dose of noise per arrow. The dotted arrow is the shortcut: pick any step , look up , draw one noise sample, and land exactly where the chain would have taken you (in distribution: the same spread of possible values). Training uses only the shortcut.

The noise schedule is the list of . This lesson uses steps with rising in a straight line from 0.0001 to 0.1: tiny doses first, while there is fine detail to lose, bigger ones once little is left.

Figure 2 · Drawn from the lesson's code

t = 0 signal √ᾱ = 1.00 t = 10 signal √ᾱ = 0.98 t = 30 signal √ᾱ = 0.80 t = 50 signal √ᾱ = 0.53 t = 100 signal √ᾱ = 0.07 The forward process: four blobs dissolve into plain noise

Five panels: the four coloured blobs at step 0 blur at step 30, merge at step 50, and are a single shapeless cloud by step 100

Reading it: each panel is the whole dataset, noised to the step in its title, with every dot keeping its blob's colour. Step 10 barely differs from the clean data (its signal share is still 0.98). By step 30 the blobs have swollen into each other, though each colour still keeps to its own corner. By step 50 the colours overlap. At step 100 only 7% of the signal is left and the colours are thoroughly mixed: a round cloud of plain noise, the same cloud whatever data you started from. That last fact is what makes generation possible: we know how to draw from that cloud.

Figure 3 · Drawn from the lesson's code

0 20 40 60 80 100 step t 0.0 0.2 0.4 0.6 0.8 1.0 share The noise schedule (T = 100, β from 0.0001 to 0.1) t = 37: half and half t = 30: 0.8 and 0.6 signal kept √ᾱ_t noise mixed in √(1 − ᾱ_t)

Two curves over 100 steps: the signal share falls from 1 to 0.07 while the noise share rises from 0 to nearly 1, crossing at step 37

Reading it: the x-axis is the step. The blue curve is the signal share and the red curve is the noise share . The two dots at step 30 are the worked example's 0.8 and 0.6. The dashed line at step 37 is where signal and noise are equal. Early steps change little (the curves are flat at the left), so the network gets plenty of practice on nearly clean data, where the fine detail lives.

In code: linear_schedule builds the list of as a NoiseSchedule, whose NoiseSchedule.alpha_bars is the running product; noise_step is one step and add_noise is the shortcut.

Why it matters in practice. The forward process has no parameters and needs no training. Its only job is to manufacture unlimited practice material, and thanks to the shortcut, any clean example can be turned into a practice question at any noise level in one line.

Chapter 2

Step 2: learning to denoise

Everyday picture A photo restorer wants to practise removing specks of dust. They take clean photos and sprinkle dust on them themselves, keeping a note of exactly where every speck went. Now they can practise all day: guess where the dust is, then check against the note. The homework is unlimited and every answer is known.

Tiny worked example Take the pixel from before: , noise , step 30, so . The network is shown only and the step number 30, and asked: "what noise was added?" The right answer is 0.5. If it guesses 0.4, it is off by 0.1 and scores . Lower is better.

Why guess the noise rather than the clean value? The two are equivalent: given , and the noise, the clean value follows by undoing the shortcut, . Guessing the noise simply works better in practice, because the target always has the same size (a standard-normal draw), whatever the step.

Figure 6 · Diagram

Reading it: three random choices feed in from the left: which example, which step, which noise. The shortcut mixes them into a noisy point. The network sees only the noisy point and the step (not the noise, and not the clean point). The noise takes the lower path straight to the loss, where the guess is graded against it. Everything after the loss is ordinary training, exactly as in primer.ml.neural_net.
Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
"theta": all of the network's weights, the knobs training turns about 5,000 numbers here
the network's guess of the noise, given the noisy point and the step; written ("epsilon hat") (0.4, −0.8)
the noise that was really added (a 2-number vector in our world) (0.5, −1.0)
the error: how far off the guess is, per number (0.1, −0.2)
squared length: square every entry and add them up
expectation: the average over many random choices of example, step and noise; in code, the average over a batch the batch average
the loss: one number saying how bad the guesses are with these weights about 0.6 after training

In words: "over many random examples, steps and noises, average how far the network's noise guess is from the real noise, measured as squared distance."

With the numbers: true noise (0.5, −1.0), guess (0.4, −0.8): error (0.1, −0.2), squared length . A network that always guesses zero scores the average of , which is 1 + 1 = 2 (each standard-normal number has variance 1). That is the score to beat.

Level 3: in Python
import random
eps = [0.5, -1.0]
guess = [0.4, -0.8]
# ‖ε − ε̂‖²: square each error and add
round(sum((e - g) ** 2 for e, g in zip(eps, guess)), 2)  # → 0.05
# the baseline: always guessing 0, averaged over many draws of 2-D noise
random.seed(0)
draws = [random.gauss(0, 1) ** 2 + random.gauss(0, 1) ** 2 for _ in range(100_000)]
round(sum(draws) / len(draws), 1)  # → 2.0

This is plain regression (predicting numbers, graded by squared error), the most stable kind of training there is. There is no opponent to balance and no trick: that is a big part of why diffusion displaced GANs.

The network. Our denoiser is a small multi-layer network: the noisy point (2 numbers), plus the step turned into 8 numbers of sines and cosines (a single number "30" is hard for a small network to use; a spread of waves at different speeds is easy, the same trick as the sinusoidal positions in primer.ml.positional), go through two hidden layers of 64 ReLU units and come out as 2 numbers, the noise guess. It trains for 1,500 steps of Adam (primer.ml.optimizers) on batches of 256, in about half a second.

Figure 5 · Drawn from the lesson's code

0 200 400 600 800 1000 1200 1400 training step 0.0 0.5 1.0 1.5 2.0 2.5 ‖ε − ε̂‖², batch average Learning to guess the noise guessing ε = 0 scores 2 each batch average of 25 batches

A training curve that starts at 2, the score of guessing zero, drops below 0.7 within a hundred steps and settles near 0.6

Reading it: the x-axis is training steps and the y-axis is the batch's average squared error. The dashed red line at 2 is the "always guess zero" baseline, and the untrained network starts right on it (its last layer starts near zero). The loss plunges within a hundred steps, then creeps down. It never reaches 0, and cannot: at low noise levels the noise is hidden inside the blob's own natural spread, so even a perfect network can only guess its average. The floor near 0.6 is that honest uncertainty, not a failure.

What the noise guess really is: the score

Flip the noise guess around and it becomes an arrow pointing toward the data. If the noise pushed the point up and to the right, the way back is down and to the left. That arrow has a name, the score: at any point, the direction in which the data gets more crowded, fastest. Learning to guess noise is secretly learning the score at every noise level, which is why diffusion models are also called score-based models, and why the training trick is called denoising score matching.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
how crowded the noisy data is at point at step (its probability density) the standard bell curve
the logarithm of that density; crowdedness on a scale where multiplying becomes adding + a constant
"grad": the slope, in every direction, as moves; for one number, just the slope slope of is
the score: which way is uphill in crowdedness at
the network's noise guess 0.9
the noise share, which converts "noise" units into "distance" units 0.6
"approximately equal": exact for a perfect guesser

In words: "the direction toward the data is the noise guess, flipped and divided by the noise share."

With the numbers: take data that is already standard-normal, so every noisy version is the plain bell curve, whose score at is (the slope of ). At with , the best possible noise guess is , and the formula gives : exactly the score, pointing back toward the crowd at 0.

Level 3: in Python
import math
x_t, alpha_bar = 1.5, 0.64
# the best noise guess for standard-normal data: √(1 − ᾱ) · x_t
eps_hat = math.sqrt(1 - alpha_bar) * x_t
round(eps_hat, 2)  # → 0.9
# −ε̂ / √(1 − ᾱ): the score, which for the bell curve is −x
round(-eps_hat / math.sqrt(1 - alpha_bar), 2)  # → -1.5

In code: DenoiserMLP is the network with its hand-written backward pass (denoiser_gradient_check compares it with finite differences), time_features turns the step into waves, train_noise_predictor is the training loop in the diagram and returns a NoisePredictor, and noise_to_score is the score formula.

Why it matters in practice. This loss, a squared error on guessed noise, is the training loss of the original DDPM paper and of the first Stable Diffusion models. The network is far bigger and the data is images, but the training loop is the boxes above.

Chapter 3

Step 3: sampling, running the process backwards

Everyday picture Back to the sculptor. Look at the block, guess which bits are "not horse", chip away a small part of them, look again. Never chip away everything you think is wrong at once: the guess is rough when the block is rough, and it gets better as the shape emerges.

Tiny worked example A point sits at at a step with and . The network guesses the noise in it is . One step back removes a scaled share of that guess, undoes the shrinking, and (except on the very last step) adds a little fresh noise.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the current, noisier point 1.0
the network's noise guess, 0.5
this step's noise dose, from the schedule 0.19
how much of the guess belongs to this one step: only a sliver of all the noise was added here
undo this step's shrink
fresh standard-normal noise; 0 on the final step 0 or 0.5
a small random wobble
the slightly cleaner point 0.9, or 1.118 with the wobble

In words: "subtract this step's share of the guessed noise, scale back up, and add a small fresh wobble."

With the numbers: ; with the wobble : .

Level 3: in Python
import math
x_t, eps_hat = 1.0, 0.5
beta, alpha_bar = 0.19, 0.75
# remove this step's share of the guessed noise, then undo the shrink
mean = (x_t - beta / math.sqrt(1 - alpha_bar) * eps_hat) / math.sqrt(1 - beta)
round(mean, 4)  # → 0.9
# add a small fresh wobble z
z = 0.5
round(mean + math.sqrt(beta) * z, 4)  # → 1.1179

Why add noise while trying to remove it? The network's guess is an average over every clean point that could have produced . Stepping to that average would drift every sample toward the safe, blurry middle. The fresh wobble keeps each sample committed to one specific possibility, so the samples spread out over the whole data, as the real data does. This recipe is DDPM (denoising diffusion probabilistic models).

Figure 8 · Diagram

Reading it: the loop runs once per step, times here and 1,000 in the original paper. Each lap costs one full run of the network, so the number of laps is the price of a sample. Keep that in mind: the rest of the lesson is largely about making that loop shorter.

Figure 7 · Drawn from the lesson's code

t = 100 t = 40 t = 20 t = 10 t = 0 Sampling: 100 small denoising steps turn noise (purple) into the blobs (grey)

Five panels of purple dots over grey blobs: a round noise cloud at step 100 is still shapeless at 40 and 20, splits into four clusters by 10, and matches the blobs at 0

Reading it: grey dots are the real data, purple dots are 600 samples being generated, shown at five moments of the backwards run. For the first half almost nothing seems to happen: the network makes coarse, whole-cloud decisions while the noise is still loud. Between steps 20 and 10 the cloud splits toward the four corners, and the last few steps tighten each group onto its blob. Nobody told the network there were four blobs; it learned that from guessing noise.

Fewer steps: jump, don't crawl (DDIM)

Everyday picture A sculptor in a hurry doesn't chip a sliver per look. They glance at the block, picture the finished horse, then rough the block out to a state only slightly less finished than that, and look again. Big confident moves, a handful of looks.

Tiny worked example From the Step 1 pixel: at , and suppose the network guesses the noise perfectly, . First predict the finished value, then re-noise it to a much quieter level, , reusing the same noise guess instead of drawing new noise.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the current point, at step 1.9
the network's noise guess 0.5
"x-zero hat": the predicted clean point, from undoing the shortcut 2.0
the earlier step to jump to; any step before , not just 96% signal
the signal left at steps and 0.64 and 0.96
the point after the jump: the predicted clean point, re-noised lightly 2.06

In words: "guess the finished point by undoing the shortcut, then walk back to a quieter noise level along the same noise direction, with no fresh randomness."

With the numbers: $\hat x_0 = (1.9 - 0.6 \cdot 0.5) / 0.8 = 1.6 / 0.8 = 2.0x_s = \sqrt{0.96} \cdot 2.0 + \sqrt{0.04} \cdot 0.5 = 1.960 + 0.1 = 2.06\bar\alpha_s = 1\hat x_0$ itself.

Level 3: in Python
import math
x_t, eps_hat = 1.9, 0.5
alpha_bar_t, alpha_bar_s = 0.64, 0.96
# x̂₀: undo the shortcut to predict the clean point
x0_hat = (x_t - math.sqrt(1 - alpha_bar_t) * eps_hat) / math.sqrt(alpha_bar_t)
round(x0_hat, 4)  # → 2.0
# x_s: re-noise it lightly, along the same noise direction
round(math.sqrt(alpha_bar_s) * x0_hat + math.sqrt(1 - alpha_bar_s) * eps_hat, 4)  # → 2.0596

This is DDIM (denoising diffusion implicit models). It uses the very same trained network; only the sampling loop changes. Because it adds no fresh noise, the same starting noise always gives the same sample, and it can take 20 jumps instead of 100 small steps. On the blobs, 20 DDIM jumps land as close to the data as 100 DDPM steps (the spec checks both).

In code: ddpm_step and ddpm_sample are the small-step sampler (ddpm_trajectory keeps snapshots for the figure); ddim_step and ddim_sample are the jumping one; two_way_distance measures how close samples land to the data, in both directions, so piling every sample onto one blob can't score well.

Why it matters in practice. Sampling cost is the number of network calls. The original DDPM used 1,000; samplers like DDIM brought that to about 20 to 50, which is what made diffusion usable in products. Pushing toward a handful of steps is the next idea.

Chapter 4

Step 4: flow matching, the straight road

Everyday picture Diffusion's road from noise to data is a winding mountain path set by the schedule. Flow matching asks: why not draw a straight line from each noise point to a data point, and learn the current that carries you along it? The network becomes a map of arrows, like a weather map of wind: at every place and time, it says which way to drift, and how fast. To generate, drop a leaf (a noise point) on the map and let it drift.

One warning about labels. In flow-matching papers, and in this section, is the noise and is the data, and time runs from 0 (noise) to 1 (data). That is the opposite of the diffusion sections above, where was the clean data.

Tiny worked example Noise point , data point . The straight line between them passes, a quarter of the way along, through . The speed along the line is constant: it covers the distance in one unit of time, so the velocity is 3 everywhere on it. The network is trained to output 3 when shown the point at time 0.25.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
a noise point (standard normal) −1
a real data point, paired with the noise at random 2
time along the line, from 0 (noise) to 1 (data), drawn at random in training 0.25
the point on the straight line at time −0.25
the velocity of that line: where to go, and how fast; the same at every 3
the network's guessed velocity at that place and time say 2.5
, squared length and average over many random draws, as in Step 2

In words: "pick a noise point, a data point and a time; find the point that far along the straight line between them; train the network to output the line's direction and speed there."

With the numbers: ; target velocity ; a guess of 2.5 scores .

Level 3: in Python
x0, x1, t = -1.0, 2.0, 0.25
# (1 − t) x₀ + t x₁: the point on the straight line
x_t = (1 - t) * x0 + t * x1
x_t  # → -0.25
# x₁ − x₀: the velocity to learn
x1 - x0  # → 3.0
# a guess of 2.5 is graded by squared error
(2.5 - (x1 - x0)) ** 2  # → 0.25

To generate, start from noise at and follow the arrows in a few Euler steps (the simplest way to follow a velocity: move in a straight line for a short time, then look again):

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the step size in time; equal steps means 0.75
the velocity the network reports here and now 3
where you are after moving for time 2.0

In words: "move for a short time in the direction, and at the speed, the network says."

With the numbers: from , one step of at velocity 3 gives : exactly the data point. On a straight line, one step of any size is exact, because there is no bend to cut.

Level 3: in Python
x_t, v, h = -0.25, 3.0, 0.75
# x + h v: one Euler step
x_t + h * v  # → 2.0

Figure 11 · Diagram

Reading it: training (top) is the same shape as diffusion's: random choices in, a point in between, a squared-error guess. The difference is the target: a velocity along a straight line rather than hidden noise. Sampling (bottom) is a plain loop of Euler steps from to , with no schedule, no shrink factors and no fresh noise.

There is a catch, and it is instructive. Many different straight lines pass through the same point (from different noise to different data), so the network cannot know which one it is on. It learns the average of their velocities. The learned paths are therefore only roughly straight: they bend, because two paths can never cross (at any point the network gives one velocity, so two paths that met would continue together).

Figure 9 · Drawn from the lesson's code

Training: random pairs, straight lines (they cross) noise x₀ data x₁ Sampling: the learned flow (paths never cross, so they bend)

Left: 16 straight lines from noise to data that cross each other. Right: the learned flow's 16 paths from the same noise, which never cross and curve to reach the blobs

Reading it: hollow circles are noise, filled circles are where each path ends. On the left are training pairs: every line is straight, but they criss-cross because noise and data were paired at random. On the right, the learned flow starts from the same noise. Its paths never cross, so they bend to share out the blobs, mostly near the end. The straighter the paths, the fewer Euler steps you need. That is why rectified flow retrains on the model's own (noise, sample) pairs, which no longer cross: each round makes the paths straighter, until one or two steps are enough.

The average also explains the extreme case. With a single Euler step from , every noise point moves by the average velocity toward "the data in general" and lands on the data's average: the empty middle of the four blobs (the spec checks it). Straight-line training does not make one step free; straight learned paths do.

Figure 10 · Drawn from the lesson's code

1 2 3 5 10 20 50 network calls per sample 1 0 − 1 two-way distance to the data Straighter paths need fewer steps a fresh draw of the real blobs plain noise diffusion (DDIM jumps) flow matching (Euler steps)

Sample quality against the number of network calls, on log axes: at one or two calls both are about as far off as plain noise; flow matching nears the fresh-data floor by about 5 calls, diffusion by about 10

Reading it: the x-axis is network calls per sample and the y-axis is the two-way distance between samples and data (lower is better). The green floor is a fresh draw of the real blobs, as good as samples can get; the grey line is plain noise. At one or two calls neither method does much better than noise. Flow matching (teal) gets close to the floor within about 5 calls, while diffusion's DDIM jumps (purple) need about 10 to get as close. Same network size, same training time: the straighter paths need roughly half the steps.

In code: flow_point and flow_target build the training pairs, train_velocity_predictor trains a VelocityPredictor, euler_step and euler_sample follow it, and step_sweep measures both methods at each step count.

Why it matters in practice. Flow matching and rectified flow give a simpler recipe (no noise schedule to tune, just straight lines) and need fewer sampling steps. Stable Diffusion 3 is trained as a rectified flow, and many recent image and video generators follow. Under the hood it is the same family as diffusion: a network trained by squared error to point from noise toward data, followed step by step.

Chapter 5

Step 5: conditioning and guidance, generating what you asked for

Everyday picture A caricaturist draws a famous face by noticing what makes it different from an average face (the big chin, the eyebrows) and exaggerating exactly that difference. Guidance does the same: compare "what I'd draw if told what to draw" with "what I'd draw anyway", and push further in the direction of the difference.

Conditioning first: to ask for a specific corner, feed the label to the network as an extra input. Our labels are four corners, given as a one-hot vector (all zeros except a 1 in the chosen slot), plus a fifth slot that means "no label". A text-to-image model does the same with a sentence in place of a corner name.

Classifier-free guidance trains one network to answer both questions: in training, the label is hidden (replaced by "no label") on a random 20% of examples. At sampling time the network is asked twice per step, once with the label and once without, and the two guesses are mixed.

Tiny worked example At some point during sampling, the guess without the label is and the guess with the label "south-west" is . The label moves the guess by . With guidance weight , move three times as far.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
the noise guess with the label hidden (, "empty set", stands for "no label") 0.2
the noise guess given the label 0.5
what the label changes: the direction "more like " 0.3
the guidance weight (or guidance scale): 0 ignores the label, 1 is the plain labelled guess, above 1 exaggerates 3
"epsilon tilde": the guided guess, used in place of by the sampler 1.1

In words: "start from the unlabelled guess and move times as far as the label would move it."

With the numbers: . With it is 0.2 (the label is ignored); with it is 0.5 (exactly the labelled guess).

Level 3: in Python
eps_uncond, eps_cond = 0.2, 0.5
# ε̃ = ε̂_∅ + w (ε̂_c − ε̂_∅), for w = 0, 1 and 3
[round(eps_uncond + w * (eps_cond - eps_uncond), 2) for w in (0, 1, 3)]  # → [0.2, 0.5, 1.1]

Figure 13 · Diagram

Reading it: each step runs the same network twice, once with the label hidden and once with it shown, which doubles the cost of a step. The two guesses meet in the mixing box, and only the mixed guess reaches the sampler, which is otherwise unchanged. Nothing new was trained for guidance: it is a choice made at sampling time.

Figure 12 · Drawn from the lesson's code

w = 0: 27% south-west spread 0.72 w = 1: 100% south-west spread 0.27 w = 3: 100% south-west spread 0.14 Classifier-free guidance, asking for "south-west"

Three panels of 400 green samples asked to be south-west: at w = 0 they spread over all four blobs (27% south-west); at w = 1 all land on the south-west blob with spread 0.27; at w = 3 they bunch into a tight knot at its outer edge

Reading it: grey dots are the data, green dots are samples asked for "south-west" at three guidance weights. At the label is ignored, so samples land in all four corners, only about a quarter in the south-west. At nearly all land in the right blob with roughly the real blob's spread. At they are all on target but bunched into a tight knot, pushed to the side of the blob farthest from the other blobs: the most unmistakably south-west spot. That is guidance in one picture: more on-label, less varied, and exaggerated.

In code: network_inputs appends the one-hot label with its "no label" slot, train_noise_predictor hides labels at random when given them, guided_noise is the formula, ddim_sample applies it at every jump when given a label, and guidance_sweep measures the share on target and the spread at each weight.

Why it matters in practice. Text-to-image systems use guidance weights well above 1, because unguided samples follow the prompt only loosely. Turn it too high and images become oversaturated and samey, the image version of the tight knot above. The "guidance scale" slider in image tools is this .

Chapter 6

Step 6: scaling up to real images, video and audio

Everyday picture An architect doesn't design a skyscraper by placing every brick. They draw a floor plan, small enough to think about, and a builder turns the plan into a building. Latent diffusion does the creative work on a small "plan" of the image, and a separate decoder builds the pixels.

Tiny worked example A 512 × 512 colour image is $512 \times 512 \times 3 = 786{,}432$ numbers. An autoencoder (a network that squeezes an image into a small code and rebuilds it; see primer.ml.generative.autoencoders) shrinks each side by 8 and keeps 4 numbers per position: $64 \times 64 \times 4 = 16{,}384$ numbers. The denoiser now works on 48 times fewer numbers, and every one of its many steps is that much cheaper.

Level 3: the formula and its symbols

Symbols

Symbol Meaning here In the example
image height and width in pixels 512, 512
3 colour channels: red, green, blue 3
how much the autoencoder shrinks each side 8
numbers kept per latent position (latent channels) 4
shrink factor how many times fewer numbers the denoiser handles 48

In words: "count the numbers in the image, count the numbers in its latent code, and divide."

With the numbers: . The latent is then cut into 2 × 2 patches, each patch becoming one token for a transformer: tokens. A 16-frame video clip in the same latent space is 16 times that, 16,384 tokens, and since attention compares every token with every other (primer.ml.attention), it costs times as much.

Level 3: in Python
H, W, f, c = 512, 512, 8, 4
pixels = H * W * 3
latents = (H // f) * (W // f) * c
pixels, latents  # → (786432, 16384)
# shrink factor
pixels / latents  # → 48.0
# 2 × 2 patches of the 64 × 64 latent: the transformer's tokens
tokens = (64 // 2) * (64 // 2)
tokens  # → 1024
# 16 video frames: 16× the tokens, 256× the attention pairs
(16 * tokens) ** 2 // tokens ** 2  # → 256

Figure 14 · Diagram

Reading it: follow the noise from the left. It is not an image but a small latent, and all the sampling steps (the loop on the denoiser box) happen in that cheap space. The prompt takes the upper path: a text encoder, such as CLIP's text half (primer.ml.embeddings.contrastive), turns it into token vectors that the denoiser reads through cross-attention (attention whose queries come from the image patches and whose keys and values come from the text tokens). Guidance's "no label" answer is simply an empty prompt. Only at the very end does the decoder turn the latent into pixels, once.

Three more scale-ups follow the same pattern:

  • The denoiser became a transformer. Early systems used a U-Net (a convolutional network, primer.ml.cnn_rnn); the diffusion transformer (DiT) cuts the latent into patches and treats them as tokens, so the scaling lessons of language models carry over.
  • Video adds a time axis: the latent is a stack of frames, and patches span space and time. It is the same denoising, with far more tokens.
  • Audio is denoised as a spectrogram (a picture of sound: time across, pitch up, loudness as brightness) or as an audio autoencoder's latent.

Compared with the other generators in this part:

GAN (primer.ml.generative.gans) VAE (primer.ml.generative.autoencoders) Diffusion and flow matching
Training a generator against a critic: unstable reconstruct plus stay near a simple code: stable squared error on noise or velocity: stable
Samples sharp often blurry sharp
Covers all of the data? can mode collapse (ignore whole regions) yes yes
Cost of one sample one network call one network call 5 to 50 network calls

In code: latent_shrink counts pixels against latent numbers and patch_tokens counts a diffusion transformer's tokens, per frame.

Why it matters in practice. Latent space made high-resolution diffusion affordable on a single GPU, transformers made it scale, text encoders made it follow prompts, and guidance made it follow them closely. The price that remains is many network calls per sample, which is why step-reduction (DDIM, flow matching, distillation into few-step students) is where so much engineering effort goes.

Test yourself

7 questions

Answer each one out loud or on paper before you open it. If you can explain it, you know it.

Question 1Why does the network learn to guess the noise, rather than the clean picture?Think it through, then reveal

The two carry the same information: given the noisy point, the step and the noise, the clean point follows by undoing the shortcut. Guessing the noise works better in practice because the target is always the same size (a standard-normal draw) at every step, which makes one network easy to train across all noise levels. The noise guess, flipped and rescaled, is also the score: the direction toward the data.

Question 2Training only ever takes one jump from clean to noisy. Why does generation need many steps back?Think it through, then reveal

The noise guess is an average over every clean point that could have produced the noisy one. From heavy noise that average is vague, pointing at the middle of the data, so one big step lands on a blur (or, in our blobs, the empty middle). Small steps let the guess sharpen as the sample commits to one specific region, and the network is asked again at every stage.

Question 3Why does DDPM add fresh noise at each step, when the goal is to remove noise?Think it through, then reveal

Without it, every step moves toward the network's averaged guess and samples drift toward safe, typical, blurry results. The small fresh wobble keeps each sample exploring one specific possibility, so the samples cover all of the data. DDIM drops the wobble on purpose, trading that randomness for determinism and big jumps.

Question 4What does flow matching change, and why can it get away with fewer steps?Think it through, then reveal

It replaces the noise schedule with straight lines from noise to data and trains the network to output the velocity along them. Following a velocity with Euler steps is only exact on straight paths; the learned paths are straighter than diffusion's, so fewer steps cut fewer corners. They are not perfectly straight, because paths can't cross, which is what rectified flow's retraining fixes.

Question 5What does the guidance weight do, and what goes wrong if it is too large?Think it through, then reveal

It sets how far past the labelled guess to go, along the direction from the unlabelled guess to the labelled one. 0 ignores the prompt, 1 follows it plainly, above 1 exaggerates it, so samples match the prompt more reliably. Too large and samples become samey, over-saturated caricatures of the prompt, bunched into the most extreme examples.

Question 6Why do real image generators denoise in a latent space instead of on pixels?Think it through, then reveal

Most of an image's pixel values are fine texture that a decoder can fill in. An autoencoder shrinks a 512 × 512 image 48-fold into a latent that keeps the meaningful structure, and since sampling runs the denoiser dozens of times, every step becomes that much cheaper. The decoder runs just once, at the end.

Question 7When would you choose a GAN over a diffusion model?Think it through, then reveal

When one-shot speed matters most: a GAN makes a sample in a single network call, where diffusion needs several to dozens. Diffusion wins on training stability and on covering all of the data (GANs can mode-collapse), which is why it took over image generation; distillation is now closing its speed gap.

Primary sources

The papers behind this lesson

Sohl-Dickstein et al., Deep Unsupervised Learning using Nonequilibrium Thermodynamics (2015)

The original idea: destroy data slowly with noise, and learn to reverse the destruction.

The paper ↗
Song and Ermon, Generative Modeling by Estimating Gradients of the Data Distribution (2019)

Generated samples by learning the score at many noise levels and following it.

The paper ↗
Ho, Jain and Abbeel, Denoising Diffusion Probabilistic Models (2020)

The simple "guess the noise" loss and the sampler of Step 3, with the first high-quality image results.

Read the annotated companion →The paper ↗
Song, Meng and Ermon, Denoising Diffusion Implicit Models (2020)

Deterministic sampling with big jumps, using the same trained network.

Read the annotated companion →The paper ↗
Song et al., Score-Based Generative Modeling through Stochastic Differential Equations (2020)

Showed diffusion and score-based models are one family, described as continuous time processes.

Read the annotated companion →The paper ↗
Ho and Salimans, Classifier-Free Diffusion Guidance (2022)

Guidance from one network trained with and without the label, no separate classifier needed.

Read the annotated companion →The paper ↗
Rombach et al., High-Resolution Image Synthesis with Latent Diffusion Models (2021)

Denoising in an autoencoder's latent space, with text via cross-attention: the basis of Stable Diffusion.

Read the annotated companion →The paper ↗
Peebles and Xie, Scalable Diffusion Models with Transformers (2022)

Replaced the U-Net with a transformer over latent patches, and showed it improves with scale.

Read the annotated companion →The paper ↗
Lipman et al., Flow Matching for Generative Modeling (2022)

Trained continuous flows by regressing velocities along simple paths from noise to data.

Read the annotated companion →The paper ↗
Liu, Gong and Liu, Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow (2022)

Straight-line paths, and retraining on the model's own pairs to straighten the learned flow for few-step sampling.

Read the annotated companion →The paper ↗
Esser et al., Scaling Rectified Flow Transformers for High-Resolution Image Synthesis (2024)

Rectified flow with a transformer denoiser at scale: Stable Diffusion 3.

The paper ↗

Researcher's shelf

Further reading

  • Ho, Jain and Abbeel, Denoising Diffusion Probabilistic Models (2020): https://arxiv.org/abs/2006.11239
  • Song, Meng and Ermon, Denoising Diffusion Implicit Models (2020): https://arxiv.org/abs/2010.02502
  • Ho and Salimans, Classifier-Free Diffusion Guidance (2022): https://arxiv.org/abs/2207.12598
  • Lipman et al., Flow Matching for Generative Modeling (2022): https://arxiv.org/abs/2210.02747
  • Liu, Gong and Liu, Rectified Flow (2022): https://arxiv.org/abs/2209.03003
  • Rombach et al., Latent Diffusion Models (2021): https://arxiv.org/abs/2112.10752
  • Peebles and Xie, Diffusion Transformers (2022): https://arxiv.org/abs/2212.09748
  • Lilian Weng, What are Diffusion Models?: https://lilianweng.github.io/posts/2021-07-11-diffusion-models/
  • Hugging Face, The Annotated Diffusion Model (DDPM, line by line in code): https://huggingface.co/blog/annotated-diffusion
  • Hugging Face Diffusers documentation: https://huggingface.co/docs/diffusers/index

About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit 048aeaa, so the two always agree: the explanation, the code that builds it and the tests that prove it.