The lesson in one minute
What you'll be able to explain
- Forward process: mix data with a little Gaussian noise per step until only noise is left; the shortcut jumps to any step in one go.
- Training: show the network a noised example and the step; it guesses the noise; grade by squared error. Plain, stable regression, and secretly learning the score (the direction toward the data).
- Sampling: start from pure noise and repeatedly remove a little guessed noise (DDPM, with a fresh wobble each step), or take a few big deterministic jumps (DDIM).
- Flow matching: learn the velocity along straight lines from noise to data and follow it with Euler steps; straighter paths need fewer steps.
- Guidance: ask the same network with and without the label and push past the labelled guess; more on-prompt, less varied.
- At scale: denoise an autoencoder's small latent with a transformer that reads a text encoder's tokens; video and audio are the same idea with more tokens.
Level 1
The practitioner's guide
In one sentence
A diffusion model generates by learning one modest skill, "remove a little noise from this noisy example", and running it many times starting from pure noise; flow matching is the same idea along straight paths with fewer steps; together they are the engine inside today's image, video and audio generators, and the settings you meet (steps, guidance scale, resolution, seed) are the dials of that engine.
When you need it
You need this lesson the moment you generate or edit
images, video or audio: whether you call a hosted model or run an open one,
the choices you make (which model family, how many steps, what guidance
scale, what size, which sampler) are the ones below, and the bill and the
artefacts follow from them. You don't need diffusion for text (language
models generate one token at a time), for a one-pass generator in a
real-time loop (a GAN or a distilled few-step model, see
primer.ml.generative.gans), or for a tiny domain with a handful of
factors (a VAE will do, see primer.ml.generative.autoencoders). The number
that shows the naive approach failing: ask this lesson's trained model for
a sample in a single step and every sample lands on the data's average, the
empty middle between the four blobs (mean distance from the centre 0.08);
give it 20 well-placed jumps and its samples sit as close to the data as a
fresh draw of the data itself (0.056 against 0.055).
Your options
From the least commitment to the most:
| Option | What it does | What it guarantees | What it costs | Where it lives |
|---|---|---|---|---|
| A hosted image, video or audio model | Prompt in, sample out; you set steps, guidance, size and seed | No infrastructure; the vendor's sampler and safety checks | A price per sample, and only the dials the API exposes | The vendor's API |
| An open latent-diffusion model in a pipeline | A pretrained denoiser, text encoder and autoencoder you run | Full control of sampler, steps, guidance, seed and adapters | A GPU with enough memory; defaults of 50 steps and guidance 7.5 in diffusers | Your server |
| A faster sampler on the same model | DDIM or a higher-order solver takes big deterministic jumps | The same trained weights, 10 to 50 times faster than the original 1,000 steps | Quality falls off below about 10 calls | A setting in the pipeline |
| A flow-matching or rectified-flow model | Trained to follow straight paths from noise to data | Fewer steps for the same quality (near the floor by 5 calls here, against 10 for diffusion) | A model trained that way, such as Stable Diffusion 3 | The model family you download |
| A distilled few-step model | A student trained to match the full model in 1 to 4 steps, often with an adversarial loss | Real-time sampling | Some variety and detail; a teacher and a distillation run | The fast path beside the full model |
| Your own diffusion model | This lesson's training loop on your own data | A generator for a narrow domain nobody has published | Data, a training run (half a second for the lesson's toy; hundreds of GPU-days for pixel-space image models), an autoencoder if the data is large | Your training loop |
How to choose
Start from the latency you can afford, then set the dials in this order.
- Steps first. The number of network calls is the price of a sample. Start at the pipeline default (50 in diffusers), halve it while the output holds, and reach for a flow model or a distilled one when you need fewer than about 10.
- Guidance scale next. 0 ignores the prompt (27% of samples on the asked-for blob here), 1 follows it plainly (100% on target, with the data's own spread), and above 1 exaggerates it: at 3 every sample is on target but bunched into a knot at the far edge of the blob, spread 0.14 against the real 0.27. Text-to-image models default to well above 1 (7.5 in diffusers), because unguided samples follow prompts loosely; raise it for obedience, lower it for variety.
- Resolution: generate at the size the model was trained on (Stable Diffusion 1.x was trained on 512 × 512 images), and upscale afterwards. Every doubling of side length quadruples the latent and the attention cost more than that.
- Determinism: DDIM and flow samplers add no fresh noise, so one seed gives one image, which is what makes a seed reproducible and an image editable by re-running with a changed prompt.
- Whatever you pick, judge the output on variety as well as quality. Every dial that makes samples match the prompt better makes them more alike.
What it costs
Latency is steps times the cost of one denoiser call, and guidance doubles the calls per step (the network runs once with the prompt and once without). Working in an autoencoder's latent cuts each call 48-fold for a 512 × 512 image (786,432 numbers down to 16,384), which is what made high-resolution diffusion affordable on one GPU. Video multiplies it back: 16 frames of the same latent are 16 times the tokens and, because attention compares every token with every other, 256 times the attention work. The denoiser itself is now usually a transformer over latent patches, and the DiT paper found that more compute per sample (a deeper or wider model, or more tokens) gives lower FID, down to 2.27 on 256 × 256 ImageNet. Training a real model is the expensive part: pixel-space diffusion "often consumes hundreds of GPU days" in the latent-diffusion paper's own words, which is why the pretrained autoencoder and the latent exist. The training loss itself is plain squared-error regression with no adversary, which is why the runs are stable and why diffusion displaced GANs.
What breaks
- Too few steps. Below roughly 10 calls a diffusion sampler's output drifts toward the average (the empty middle here; blur or mush in images). Use a flow or distilled model instead of starving a diffusion one.
- Guidance too high. Oversaturated, samey, exaggerated images, the image version of the tight knot at weight 3; the diffusers documentation puts it as prompt adherence "at the expense of lower image quality". Pull it back toward the default.
- Guidance too low. The prompt is followed loosely or not at all, as at weight 0. A guidance scale of 1 or below switches guidance off.
- The wrong size. A model asked for a resolution or aspect ratio far from its training size composes badly (repeated subjects, stretched scenes). Generate at the native size and upscale.
- A mismatched autoencoder or scale. The latent must be decoded by the
autoencoder the denoiser was trained with, with the library's scaling
factor applied both ways, or the output is junk; see
primer.ml.generative.autoencoders. - Averaging without the wobble. DDPM adds a small fresh noise each step so samples commit to one possibility; a sampler that only ever steps to the network's average drifts to the safe, blurry middle. Deterministic samplers avoid this by jumping along the predicted noise direction, not to the mean.
- Cost that scales with frames. A video request costs attention quadratically in its token count; a short clip at a modest size is many images' worth of compute, and the bill follows.
In the wild
Stable Diffusion is the open reference: an 860-million parameter U-Net denoising a 64 × 64 × 4 latent, conditioned through cross-attention on a frozen CLIP ViT-L/14 text encoder, trained on 512 × 512 images, and run through Hugging Face diffusers with 50 steps and guidance 7.5 by default. DDIM (Song, Meng and Ermon) gave the 10 to 50 times faster deterministic sampler every pipeline offers; classifier-free guidance (Ho and Salimans) is the guidance-scale slider, trading variety for fidelity with no separate classifier; DiT (Peebles and Xie) replaced the U-Net with a transformer over latent patches and showed it scales; Stable Diffusion 3 (Esser et al.) trains a transformer as a rectified flow with 16-channel latents; Adversarial Diffusion Distillation turns a foundation model into a one-to-four-step sampler. Video generators run the same denoiser over patches that span space and time, and audio generators denoise a spectrogram or an audio autoencoder's latent. Every paper is linked at the end of the lesson.
Go deeper
Level 2 builds the whole engine on four blobs of dots: the forward process and its one-jump shortcut, the noise-guessing loss and why it is secretly learning the direction toward the data, DDPM and DDIM sampling step by step, flow matching along straight lines, guidance with its knot, and the latent-and-patches arithmetic behind real systems. If you only needed to set the dials, you are done.
Level 2
How it works, from scratch
Imagine a sculptor who cannot carve a statue in one go. Hand them a shapeless block and ask for a horse, and they freeze. But hand them a slightly rough horse, and they can always make it a little less rough: smooth this bump, deepen that line. Now chain that one modest skill. Start from a shapeless block, make it a bit more horse-like, then a bit more, a hundred times over, and a horse comes out.
A diffusion model is that sculptor. It never learns "draw a picture from nothing", which is very hard. It learns "here is a picture with a little too much noise in it: take some of the noise away", which is much easier. And the practice material is free: take any real picture, add noise yourself, and you know exactly what the answer should be.
Figure 1 · Diagram
flowchart LR D["Real data<br/>(a photo)"] -- "add a little noise,<br/>T times (fixed, no learning)" --> N["Pure noise<br/>(TV static)"] N -- "remove a little noise,<br/>T times (learned)" --> S["A new sample<br/>(a photo no one took)"]
The tiny world this lesson works in. A real image is a long list of numbers, one per pixel colour. To keep everything small enough to see, our "images" have just two numbers, so each one is a dot on a flat plane. The real data is four round blobs of dots, one in each corner (north-east, north-west, south-west, south-east), 600 dots in all. The job: learn to make new dots that land in the blobs, starting from random noise. Every idea below works unchanged on a 512 × 512 colour image; it just has 786,432 numbers per "dot" instead of 2.
Notice one thing about the blobs now, because it matters later: their average is the empty middle. A model that plays safe and outputs "the average answer" puts its dots where no real data lives. The same failure makes blurry images: the average of many sharp faces is a blur.
Chapter 1
Step 1: the forward process, adding noise a little at a time
Everyday picture Think of an old television slowly losing its signal. Turn the static dial up one notch: the picture fades a little and a little snow creeps in. Turn it a hundred notches and you see only snow, and nothing in it tells you what the programme was.
Tiny worked example Follow one pixel whose clean value is . Each step keeps most of the signal and mixes in a little fresh random noise. The noise is a random number drawn from a standard normal distribution: the bell curve centred on 0 whose typical size is 1 (values like 0.5, −1.2 and 0.1 are common, 3 is rare). Say step 1 mixes in a share of noise, and the random draw is .
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the step number, from 1 to (the last step) | 1 | |
| the value before this step ( is the clean data) | 2.0 | |
| the value after this step | 2.06 | |
| "beta": how much noise this step mixes in, a number between 0 and 1 | 0.1 | |
| how much of the old value survives: it shrinks a little | ||
| "epsilon": fresh noise drawn from the standard normal bell curve | 0.5 | |
| how much noise is mixed in |
In words: "shrink the old value a little, then add a little random noise."
With the numbers: $x_1 = \sqrt{0.9} \cdot 2.0 + \sqrt{0.1} \cdot 0.5 = 1.897 + 0.158 = 2.06$.
Level 3: in Python
import math
x_prev, beta, eps = 2.0, 0.1, 0.5
# √(1 − β) x_{t−1}: shrink the old value
round(math.sqrt(1 - beta) * x_prev, 3) # → 1.897
# √β ε: the noise mixed in
round(math.sqrt(beta) * eps, 3) # → 0.158
round(math.sqrt(1 - beta) * x_prev + math.sqrt(beta) * eps, 2) # → 2.06
Why the square roots? A value's variance (its typical squared size) is what matters, and variances of independent things add. If the old value has variance 1, the new one has variance : exactly 1 again. The square roots keep every step at the same overall size, so after many steps the value is noise of size 1, not a number that has blown up or faded to zero.
The shortcut: jump straight to any step
Running steps one by one is slow, and training needs noisy examples at every step. Luckily, adding noise twice is the same as adding it once in a bigger dose, so there is a formula that jumps straight to step . First, multiply up how much signal survives every step so far:
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the share of variance step keeps (often written ) | 0.9, then 0.8 | |
"multiply together, for = 1, 2, …, " (see primer.notation) |
||
| "alpha bar": the share of the original signal's variance left after steps; 1 means clean, 0 means pure noise | 0.72; or 0.64 in the jump below | |
| the clean data | 2.0 | |
| one fresh standard-normal draw, standing in for all the noise added so far | 0.5 | |
| the signal share: how much of is still in | ||
| the noise share |
In words: "multiply together the survival share of every step so far; then the noisy value is that much of the clean value plus the rest in noise."
With the numbers: two steps with then leave . In this lesson's schedule, step 30 leaves , so the pixel at 2.0 with noise 0.5 lands at .
Level 3: in Python
import math
betas = [0.1, 0.2]
alpha_bar = 1.0
# ∏ (1 − β_s): multiply up what each step keeps
for beta in betas:
alpha_bar *= 1 - beta
round(alpha_bar, 2) # → 0.72
# the jump to step 30, where ᾱ = 0.64
x0, eps, alpha_bar_30 = 2.0, 0.5, 0.64
round(math.sqrt(alpha_bar_30) * x0 + math.sqrt(1 - alpha_bar_30) * eps, 2) # → 1.9
Figure 4 · Diagram
flowchart LR X0["x₀<br/>clean"] --> X1["x₁"] --> X2["x₂"] --> DOTS["…"] --> XT["x_T<br/>pure noise"] X0 -. "shortcut: √ᾱ_t · x₀ + √(1 − ᾱ_t) · ε" .-> XM["x_t<br/>any step, in one jump"]
The noise schedule is the list of . This lesson uses steps with rising in a straight line from 0.0001 to 0.1: tiny doses first, while there is fine detail to lose, bigger ones once little is left.
Figure 2 · Drawn from the lesson's code
Five panels: the four coloured blobs at step 0 blur at step 30, merge at step 50, and are a single shapeless cloud by step 100
Figure 3 · Drawn from the lesson's code
Two curves over 100 steps: the signal share falls from 1 to 0.07 while the noise share rises from 0 to nearly 1, crossing at step 37
In code: linear_schedule builds the list of as a NoiseSchedule, whose NoiseSchedule.alpha_bars is the running product; noise_step is one step and add_noise is the shortcut.
Why it matters in practice. The forward process has no parameters and needs no training. Its only job is to manufacture unlimited practice material, and thanks to the shortcut, any clean example can be turned into a practice question at any noise level in one line.
Chapter 2
Step 2: learning to denoise
Everyday picture A photo restorer wants to practise removing specks of dust. They take clean photos and sprinkle dust on them themselves, keeping a note of exactly where every speck went. Now they can practise all day: guess where the dust is, then check against the note. The homework is unlimited and every answer is known.
Tiny worked example Take the pixel from before: , noise , step 30, so . The network is shown only and the step number 30, and asked: "what noise was added?" The right answer is 0.5. If it guesses 0.4, it is off by 0.1 and scores . Lower is better.
Why guess the noise rather than the clean value? The two are equivalent: given , and the noise, the clean value follows by undoing the shortcut, . Guessing the noise simply works better in practice, because the target always has the same size (a standard-normal draw), whatever the step.
Figure 6 · Diagram
flowchart LR D["pick a real point x₀"] --> MIX T["pick a random step t"] --> MIX E["draw noise ε"] --> MIX["mix with the shortcut<br/>x_t = √ᾱ_t x₀ + √(1 − ᾱ_t) ε"] MIX --> NET["network ε_θ<br/>sees x_t and t"] NET --> G["guess ε̂"] G --> L["loss ‖ε − ε̂‖²"] E --> L L --> U["nudge the weights<br/>(backprop + Adam)"]
primer.ml.neural_net.Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| "theta": all of the network's weights, the knobs training turns | about 5,000 numbers here | |
| the network's guess of the noise, given the noisy point and the step; written ("epsilon hat") | (0.4, −0.8) | |
| the noise that was really added (a 2-number vector in our world) | (0.5, −1.0) | |
| the error: how far off the guess is, per number | (0.1, −0.2) | |
| squared length: square every entry and add them up | ||
| expectation: the average over many random choices of example, step and noise; in code, the average over a batch | the batch average | |
| the loss: one number saying how bad the guesses are with these weights | about 0.6 after training |
In words: "over many random examples, steps and noises, average how far the network's noise guess is from the real noise, measured as squared distance."
With the numbers: true noise (0.5, −1.0), guess (0.4, −0.8): error (0.1, −0.2), squared length . A network that always guesses zero scores the average of , which is 1 + 1 = 2 (each standard-normal number has variance 1). That is the score to beat.
Level 3: in Python
import random
eps = [0.5, -1.0]
guess = [0.4, -0.8]
# ‖ε − ε̂‖²: square each error and add
round(sum((e - g) ** 2 for e, g in zip(eps, guess)), 2) # → 0.05
# the baseline: always guessing 0, averaged over many draws of 2-D noise
random.seed(0)
draws = [random.gauss(0, 1) ** 2 + random.gauss(0, 1) ** 2 for _ in range(100_000)]
round(sum(draws) / len(draws), 1) # → 2.0
This is plain regression (predicting numbers, graded by squared error), the most stable kind of training there is. There is no opponent to balance and no trick: that is a big part of why diffusion displaced GANs.
The network. Our denoiser is a small multi-layer network: the noisy
point (2 numbers), plus the step turned into 8 numbers of sines and cosines
(a single number "30" is hard for a small network to use; a spread of waves
at different speeds is easy, the same trick as the sinusoidal positions in
primer.ml.positional), go through two hidden layers of 64 ReLU units
and come out as 2 numbers, the noise guess. It trains for 1,500 steps of
Adam (primer.ml.optimizers) on batches of 256, in about half a second.
Figure 5 · Drawn from the lesson's code
A training curve that starts at 2, the score of guessing zero, drops below 0.7 within a hundred steps and settles near 0.6
What the noise guess really is: the score
Flip the noise guess around and it becomes an arrow pointing toward the data. If the noise pushed the point up and to the right, the way back is down and to the left. That arrow has a name, the score: at any point, the direction in which the data gets more crowded, fastest. Learning to guess noise is secretly learning the score at every noise level, which is why diffusion models are also called score-based models, and why the training trick is called denoising score matching.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| how crowded the noisy data is at point at step (its probability density) | the standard bell curve | |
| the logarithm of that density; crowdedness on a scale where multiplying becomes adding | + a constant | |
| "grad": the slope, in every direction, as moves; for one number, just the slope | slope of is | |
| the score: which way is uphill in crowdedness | at | |
| the network's noise guess | 0.9 | |
| the noise share, which converts "noise" units into "distance" units | 0.6 | |
| "approximately equal": exact for a perfect guesser |
In words: "the direction toward the data is the noise guess, flipped and divided by the noise share."
With the numbers: take data that is already standard-normal, so every noisy version is the plain bell curve, whose score at is (the slope of ). At with , the best possible noise guess is , and the formula gives : exactly the score, pointing back toward the crowd at 0.
Level 3: in Python
import math
x_t, alpha_bar = 1.5, 0.64
# the best noise guess for standard-normal data: √(1 − ᾱ) · x_t
eps_hat = math.sqrt(1 - alpha_bar) * x_t
round(eps_hat, 2) # → 0.9
# −ε̂ / √(1 − ᾱ): the score, which for the bell curve is −x
round(-eps_hat / math.sqrt(1 - alpha_bar), 2) # → -1.5
In code: DenoiserMLP is the network with its hand-written backward pass (denoiser_gradient_check compares it with finite differences), time_features turns the step into waves, train_noise_predictor is the training loop in the diagram and returns a NoisePredictor, and noise_to_score is the score formula.
Why it matters in practice. This loss, a squared error on guessed noise, is the training loss of the original DDPM paper and of the first Stable Diffusion models. The network is far bigger and the data is images, but the training loop is the boxes above.
Chapter 3
Step 3: sampling, running the process backwards
Everyday picture Back to the sculptor. Look at the block, guess which bits are "not horse", chip away a small part of them, look again. Never chip away everything you think is wrong at once: the guess is rough when the block is rough, and it gets better as the shape emerges.
Tiny worked example A point sits at at a step with and . The network guesses the noise in it is . One step back removes a scaled share of that guess, undoes the shrinking, and (except on the very last step) adds a little fresh noise.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the current, noisier point | 1.0 | |
| the network's noise guess, | 0.5 | |
| this step's noise dose, from the schedule | 0.19 | |
| how much of the guess belongs to this one step: only a sliver of all the noise was added here | ||
| undo this step's shrink | ||
| fresh standard-normal noise; 0 on the final step | 0 or 0.5 | |
| a small random wobble | ||
| the slightly cleaner point | 0.9, or 1.118 with the wobble |
In words: "subtract this step's share of the guessed noise, scale back up, and add a small fresh wobble."
With the numbers: ; with the wobble : .
Level 3: in Python
import math
x_t, eps_hat = 1.0, 0.5
beta, alpha_bar = 0.19, 0.75
# remove this step's share of the guessed noise, then undo the shrink
mean = (x_t - beta / math.sqrt(1 - alpha_bar) * eps_hat) / math.sqrt(1 - beta)
round(mean, 4) # → 0.9
# add a small fresh wobble z
z = 0.5
round(mean + math.sqrt(beta) * z, 4) # → 1.1179
Why add noise while trying to remove it? The network's guess is an average over every clean point that could have produced . Stepping to that average would drift every sample toward the safe, blurry middle. The fresh wobble keeps each sample committed to one specific possibility, so the samples spread out over the whole data, as the real data does. This recipe is DDPM (denoising diffusion probabilistic models).
Figure 8 · Diagram
flowchart LR
N["x_T: draw pure noise"] --> G["network guesses<br/>the noise ε̂ at step t"]
G --> R["remove this step's share,<br/>undo the shrink"]
R --> W["add a small fresh wobble<br/>(not on the last step)"]
W --> Q{"t = 0?"}
Q -- "no: t ← t − 1" --> G
Q -- "yes" --> OUT["x₀: a new sample"]
Figure 7 · Drawn from the lesson's code
Five panels of purple dots over grey blobs: a round noise cloud at step 100 is still shapeless at 40 and 20, splits into four clusters by 10, and matches the blobs at 0
Fewer steps: jump, don't crawl (DDIM)
Everyday picture A sculptor in a hurry doesn't chip a sliver per look. They glance at the block, picture the finished horse, then rough the block out to a state only slightly less finished than that, and look again. Big confident moves, a handful of looks.
Tiny worked example From the Step 1 pixel: at , and suppose the network guesses the noise perfectly, . First predict the finished value, then re-noise it to a much quieter level, , reusing the same noise guess instead of drawing new noise.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the current point, at step | 1.9 | |
| the network's noise guess | 0.5 | |
| "x-zero hat": the predicted clean point, from undoing the shortcut | 2.0 | |
| the earlier step to jump to; any step before , not just | 96% signal | |
| the signal left at steps and | 0.64 and 0.96 | |
| the point after the jump: the predicted clean point, re-noised lightly | 2.06 |
In words: "guess the finished point by undoing the shortcut, then walk back to a quieter noise level along the same noise direction, with no fresh randomness."
With the numbers: $\hat x_0 = (1.9 - 0.6 \cdot 0.5) / 0.8 = 1.6 / 0.8 = 2.0x_s = \sqrt{0.96} \cdot 2.0 + \sqrt{0.04} \cdot 0.5 = 1.960 + 0.1 = 2.06\bar\alpha_s = 1\hat x_0$ itself.
Level 3: in Python
import math
x_t, eps_hat = 1.9, 0.5
alpha_bar_t, alpha_bar_s = 0.64, 0.96
# x̂₀: undo the shortcut to predict the clean point
x0_hat = (x_t - math.sqrt(1 - alpha_bar_t) * eps_hat) / math.sqrt(alpha_bar_t)
round(x0_hat, 4) # → 2.0
# x_s: re-noise it lightly, along the same noise direction
round(math.sqrt(alpha_bar_s) * x0_hat + math.sqrt(1 - alpha_bar_s) * eps_hat, 4) # → 2.0596
This is DDIM (denoising diffusion implicit models). It uses the very same trained network; only the sampling loop changes. Because it adds no fresh noise, the same starting noise always gives the same sample, and it can take 20 jumps instead of 100 small steps. On the blobs, 20 DDIM jumps land as close to the data as 100 DDPM steps (the spec checks both).
In code: ddpm_step and ddpm_sample are the small-step sampler (ddpm_trajectory keeps snapshots for the figure); ddim_step and ddim_sample are the jumping one; two_way_distance measures how close samples land to the data, in both directions, so piling every sample onto one blob can't score well.
Why it matters in practice. Sampling cost is the number of network calls. The original DDPM used 1,000; samplers like DDIM brought that to about 20 to 50, which is what made diffusion usable in products. Pushing toward a handful of steps is the next idea.
Chapter 4
Step 4: flow matching, the straight road
Everyday picture Diffusion's road from noise to data is a winding mountain path set by the schedule. Flow matching asks: why not draw a straight line from each noise point to a data point, and learn the current that carries you along it? The network becomes a map of arrows, like a weather map of wind: at every place and time, it says which way to drift, and how fast. To generate, drop a leaf (a noise point) on the map and let it drift.
One warning about labels. In flow-matching papers, and in this section, is the noise and is the data, and time runs from 0 (noise) to 1 (data). That is the opposite of the diffusion sections above, where was the clean data.
Tiny worked example Noise point , data point . The straight line between them passes, a quarter of the way along, through . The speed along the line is constant: it covers the distance in one unit of time, so the velocity is 3 everywhere on it. The network is trained to output 3 when shown the point at time 0.25.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| a noise point (standard normal) | −1 | |
| a real data point, paired with the noise at random | 2 | |
| time along the line, from 0 (noise) to 1 (data), drawn at random in training | 0.25 | |
| the point on the straight line at time | −0.25 | |
| the velocity of that line: where to go, and how fast; the same at every | 3 | |
| the network's guessed velocity at that place and time | say 2.5 | |
| , | squared length and average over many random draws, as in Step 2 |
In words: "pick a noise point, a data point and a time; find the point that far along the straight line between them; train the network to output the line's direction and speed there."
With the numbers: ; target velocity ; a guess of 2.5 scores .
Level 3: in Python
x0, x1, t = -1.0, 2.0, 0.25
# (1 − t) x₀ + t x₁: the point on the straight line
x_t = (1 - t) * x0 + t * x1
x_t # → -0.25
# x₁ − x₀: the velocity to learn
x1 - x0 # → 3.0
# a guess of 2.5 is graded by squared error
(2.5 - (x1 - x0)) ** 2 # → 0.25
To generate, start from noise at and follow the arrows in a few Euler steps (the simplest way to follow a velocity: move in a straight line for a short time, then look again):
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the step size in time; equal steps means | 0.75 | |
| the velocity the network reports here and now | 3 | |
| where you are after moving for time | 2.0 |
In words: "move for a short time in the direction, and at the speed, the network says."
With the numbers: from , one step of at velocity 3 gives : exactly the data point. On a straight line, one step of any size is exact, because there is no bend to cut.
Level 3: in Python
x_t, v, h = -0.25, 3.0, 0.75
# x + h v: one Euler step
x_t + h * v # → 2.0
Figure 11 · Diagram
flowchart LR
subgraph TRAIN["Training"]
direction LR
P["pair noise x₀<br/>with data x₁"] --> L1["point on the line<br/>(1 − t) x₀ + t x₁"]
L1 --> V["network guesses v"]
V --> LOSS["loss ‖v − (x₁ − x₀)‖²"]
end
subgraph SAMPLE["Sampling"]
direction LR
Z["draw noise, t = 0"] --> STEP["x ← x + h · v(x, t)"]
STEP --> C{"t = 1?"}
C -- no --> STEP
C -- yes --> S["a new sample"]
end
There is a catch, and it is instructive. Many different straight lines pass through the same point (from different noise to different data), so the network cannot know which one it is on. It learns the average of their velocities. The learned paths are therefore only roughly straight: they bend, because two paths can never cross (at any point the network gives one velocity, so two paths that met would continue together).
Figure 9 · Drawn from the lesson's code
Left: 16 straight lines from noise to data that cross each other. Right: the learned flow's 16 paths from the same noise, which never cross and curve to reach the blobs
The average also explains the extreme case. With a single Euler step from , every noise point moves by the average velocity toward "the data in general" and lands on the data's average: the empty middle of the four blobs (the spec checks it). Straight-line training does not make one step free; straight learned paths do.
Figure 10 · Drawn from the lesson's code
Sample quality against the number of network calls, on log axes: at one or two calls both are about as far off as plain noise; flow matching nears the fresh-data floor by about 5 calls, diffusion by about 10
In code: flow_point and flow_target build the training pairs, train_velocity_predictor trains a VelocityPredictor, euler_step and euler_sample follow it, and step_sweep measures both methods at each step count.
Why it matters in practice. Flow matching and rectified flow give a simpler recipe (no noise schedule to tune, just straight lines) and need fewer sampling steps. Stable Diffusion 3 is trained as a rectified flow, and many recent image and video generators follow. Under the hood it is the same family as diffusion: a network trained by squared error to point from noise toward data, followed step by step.
Chapter 5
Step 5: conditioning and guidance, generating what you asked for
Everyday picture A caricaturist draws a famous face by noticing what makes it different from an average face (the big chin, the eyebrows) and exaggerating exactly that difference. Guidance does the same: compare "what I'd draw if told what to draw" with "what I'd draw anyway", and push further in the direction of the difference.
Conditioning first: to ask for a specific corner, feed the label to the network as an extra input. Our labels are four corners, given as a one-hot vector (all zeros except a 1 in the chosen slot), plus a fifth slot that means "no label". A text-to-image model does the same with a sentence in place of a corner name.
Classifier-free guidance trains one network to answer both questions: in training, the label is hidden (replaced by "no label") on a random 20% of examples. At sampling time the network is asked twice per step, once with the label and once without, and the two guesses are mixed.
Tiny worked example At some point during sampling, the guess without the label is and the guess with the label "south-west" is . The label moves the guess by . With guidance weight , move three times as far.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| the noise guess with the label hidden (, "empty set", stands for "no label") | 0.2 | |
| the noise guess given the label | 0.5 | |
| what the label changes: the direction "more like " | 0.3 | |
| the guidance weight (or guidance scale): 0 ignores the label, 1 is the plain labelled guess, above 1 exaggerates | 3 | |
| "epsilon tilde": the guided guess, used in place of by the sampler | 1.1 |
In words: "start from the unlabelled guess and move times as far as the label would move it."
With the numbers: . With it is 0.2 (the label is ignored); with it is 0.5 (exactly the labelled guess).
Level 3: in Python
eps_uncond, eps_cond = 0.2, 0.5
# ε̃ = ε̂_∅ + w (ε̂_c − ε̂_∅), for w = 0, 1 and 3
[round(eps_uncond + w * (eps_cond - eps_uncond), 2) for w in (0, 1, 3)] # → [0.2, 0.5, 1.1]
Figure 13 · Diagram
flowchart LR X["noisy point x_t, step t"] --> A["network, label hidden"] X --> B["network, label = south-west"] A --> U["ε̂_∅"] B --> C["ε̂_c"] U --> MIX["ε̃ = ε̂_∅ + w (ε̂_c − ε̂_∅)"] C --> MIX MIX --> STEP["one sampler step<br/>(DDIM or DDPM)"]
Figure 12 · Drawn from the lesson's code
Three panels of 400 green samples asked to be south-west: at w = 0 they spread over all four blobs (27% south-west); at w = 1 all land on the south-west blob with spread 0.27; at w = 3 they bunch into a tight knot at its outer edge
In code: network_inputs appends the one-hot label with its "no label" slot, train_noise_predictor hides labels at random when given them, guided_noise is the formula, ddim_sample applies it at every jump when given a label, and guidance_sweep measures the share on target and the spread at each weight.
Why it matters in practice. Text-to-image systems use guidance weights well above 1, because unguided samples follow the prompt only loosely. Turn it too high and images become oversaturated and samey, the image version of the tight knot above. The "guidance scale" slider in image tools is this .
Chapter 6
Step 6: scaling up to real images, video and audio
Everyday picture An architect doesn't design a skyscraper by placing every brick. They draw a floor plan, small enough to think about, and a builder turns the plan into a building. Latent diffusion does the creative work on a small "plan" of the image, and a separate decoder builds the pixels.
Tiny worked example A 512 × 512 colour image is $512 \times 512 \times 3
= 786{,}432$ numbers. An autoencoder (a network that squeezes an image
into a small code and rebuilds it; see primer.ml.generative.autoencoders)
shrinks each side by 8 and keeps 4 numbers per position: $64 \times 64 \times
4 = 16{,}384$ numbers. The denoiser now works on 48 times fewer numbers, and
every one of its many steps is that much cheaper.
Level 3: the formula and its symbols
Symbols
| Symbol | Meaning here | In the example |
|---|---|---|
| image height and width in pixels | 512, 512 | |
| 3 | colour channels: red, green, blue | 3 |
| how much the autoencoder shrinks each side | 8 | |
| numbers kept per latent position (latent channels) | 4 | |
| shrink factor | how many times fewer numbers the denoiser handles | 48 |
In words: "count the numbers in the image, count the numbers in its latent code, and divide."
With the numbers: . The latent is then cut
into 2 × 2 patches, each patch becoming one token for a transformer:
tokens. A 16-frame video clip in the same
latent space is 16 times that, 16,384 tokens, and since attention compares
every token with every other (primer.ml.attention), it costs
times as much.
Level 3: in Python
H, W, f, c = 512, 512, 8, 4
pixels = H * W * 3
latents = (H // f) * (W // f) * c
pixels, latents # → (786432, 16384)
# shrink factor
pixels / latents # → 48.0
# 2 × 2 patches of the 64 × 64 latent: the transformer's tokens
tokens = (64 // 2) * (64 // 2)
tokens # → 1024
# 16 video frames: 16× the tokens, 256× the attention pairs
(16 * tokens) ** 2 // tokens ** 2 # → 256
Figure 14 · Diagram
flowchart LR P["prompt: 'a red fox in snow'"] --> TE["text encoder<br/>(e.g. CLIP's text tower)"] TE --> TOK["text token vectors"] N["random latent noise<br/>64 × 64 × 4"] --> DEN TOK --> DEN["denoiser: a transformer<br/>over latent patches,<br/>attending to the text"] DEN -- "20 to 50 steps,<br/>with guidance" --> DEN DEN --> LAT["clean latent"] LAT --> DEC["autoencoder decoder"] DEC --> IMG["512 × 512 image"]
primer.ml.embeddings.contrastive),
turns it into token vectors that the denoiser reads through
cross-attention (attention whose queries come from the image patches and
whose keys and values come from the text tokens). Guidance's "no label"
answer is simply an empty prompt. Only at the very end does the decoder
turn the latent into pixels, once.Three more scale-ups follow the same pattern:
- The denoiser became a transformer. Early systems used a U-Net (a
convolutional network,
primer.ml.cnn_rnn); the diffusion transformer (DiT) cuts the latent into patches and treats them as tokens, so the scaling lessons of language models carry over. - Video adds a time axis: the latent is a stack of frames, and patches span space and time. It is the same denoising, with far more tokens.
- Audio is denoised as a spectrogram (a picture of sound: time across, pitch up, loudness as brightness) or as an audio autoencoder's latent.
Compared with the other generators in this part:
GAN (primer.ml.generative.gans) |
VAE (primer.ml.generative.autoencoders) |
Diffusion and flow matching | |
|---|---|---|---|
| Training | a generator against a critic: unstable | reconstruct plus stay near a simple code: stable | squared error on noise or velocity: stable |
| Samples | sharp | often blurry | sharp |
| Covers all of the data? | can mode collapse (ignore whole regions) | yes | yes |
| Cost of one sample | one network call | one network call | 5 to 50 network calls |
In code: latent_shrink counts pixels against latent numbers and patch_tokens counts a diffusion transformer's tokens, per frame.
Why it matters in practice. Latent space made high-resolution diffusion affordable on a single GPU, transformers made it scale, text encoders made it follow prompts, and guidance made it follow them closely. The price that remains is many network calls per sample, which is why step-reduction (DDIM, flow matching, distillation into few-step students) is where so much engineering effort goes.
Test yourself
7 questions
Answer each one out loud or on paper before you open it. If you can explain it, you know it.
Question 1Why does the network learn to guess the noise, rather than the clean picture?Think it through, then reveal
The two carry the same information: given the noisy point, the step and the noise, the clean point follows by undoing the shortcut. Guessing the noise works better in practice because the target is always the same size (a standard-normal draw) at every step, which makes one network easy to train across all noise levels. The noise guess, flipped and rescaled, is also the score: the direction toward the data.
Question 2Training only ever takes one jump from clean to noisy. Why does generation need many steps back?Think it through, then reveal
The noise guess is an average over every clean point that could have produced the noisy one. From heavy noise that average is vague, pointing at the middle of the data, so one big step lands on a blur (or, in our blobs, the empty middle). Small steps let the guess sharpen as the sample commits to one specific region, and the network is asked again at every stage.
Question 3Why does DDPM add fresh noise at each step, when the goal is to remove noise?Think it through, then reveal
Without it, every step moves toward the network's averaged guess and samples drift toward safe, typical, blurry results. The small fresh wobble keeps each sample exploring one specific possibility, so the samples cover all of the data. DDIM drops the wobble on purpose, trading that randomness for determinism and big jumps.
Question 4What does flow matching change, and why can it get away with fewer steps?Think it through, then reveal
It replaces the noise schedule with straight lines from noise to data and trains the network to output the velocity along them. Following a velocity with Euler steps is only exact on straight paths; the learned paths are straighter than diffusion's, so fewer steps cut fewer corners. They are not perfectly straight, because paths can't cross, which is what rectified flow's retraining fixes.
Question 5What does the guidance weight do, and what goes wrong if it is too large?Think it through, then reveal
It sets how far past the labelled guess to go, along the direction from the unlabelled guess to the labelled one. 0 ignores the prompt, 1 follows it plainly, above 1 exaggerates it, so samples match the prompt more reliably. Too large and samples become samey, over-saturated caricatures of the prompt, bunched into the most extreme examples.
Question 6Why do real image generators denoise in a latent space instead of on pixels?Think it through, then reveal
Most of an image's pixel values are fine texture that a decoder can fill in. An autoencoder shrinks a 512 × 512 image 48-fold into a latent that keeps the meaningful structure, and since sampling runs the denoiser dozens of times, every step becomes that much cheaper. The decoder runs just once, at the end.
Question 7When would you choose a GAN over a diffusion model?Think it through, then reveal
When one-shot speed matters most: a GAN makes a sample in a single network call, where diffusion needs several to dozens. Diffusion wins on training stability and on covering all of the data (GANs can mode-collapse), which is why it took over image generation; distillation is now closing its speed gap.
Primary sources
The papers behind this lesson
The original idea: destroy data slowly with noise, and learn to reverse the destruction.
The paper ↗Generated samples by learning the score at many noise levels and following it.
The paper ↗The simple "guess the noise" loss and the sampler of Step 3, with the first high-quality image results.
Read the annotated companion →The paper ↗Deterministic sampling with big jumps, using the same trained network.
Read the annotated companion →The paper ↗Showed diffusion and score-based models are one family, described as continuous time processes.
Read the annotated companion →The paper ↗Guidance from one network trained with and without the label, no separate classifier needed.
Read the annotated companion →The paper ↗Denoising in an autoencoder's latent space, with text via cross-attention: the basis of Stable Diffusion.
Read the annotated companion →The paper ↗Replaced the U-Net with a transformer over latent patches, and showed it improves with scale.
Read the annotated companion →The paper ↗Trained continuous flows by regressing velocities along simple paths from noise to data.
Read the annotated companion →The paper ↗Straight-line paths, and retraining on the model's own pairs to straighten the learned flow for few-step sampling.
Read the annotated companion →The paper ↗Rectified flow with a transformer denoiser at scale: Stable Diffusion 3.
The paper ↗Researcher's shelf
Further reading
- Ho, Jain and Abbeel, Denoising Diffusion Probabilistic Models (2020): https://arxiv.org/abs/2006.11239
- Song, Meng and Ermon, Denoising Diffusion Implicit Models (2020): https://arxiv.org/abs/2010.02502
- Ho and Salimans, Classifier-Free Diffusion Guidance (2022): https://arxiv.org/abs/2207.12598
- Lipman et al., Flow Matching for Generative Modeling (2022): https://arxiv.org/abs/2210.02747
- Liu, Gong and Liu, Rectified Flow (2022): https://arxiv.org/abs/2209.03003
- Rombach et al., Latent Diffusion Models (2021): https://arxiv.org/abs/2112.10752
- Peebles and Xie, Diffusion Transformers (2022): https://arxiv.org/abs/2212.09748
- Lilian Weng, What are Diffusion Models?: https://lilianweng.github.io/posts/2021-07-11-diffusion-models/
- Hugging Face, The Annotated Diffusion Model (DDPM, line by line in code): https://huggingface.co/blog/annotated-diffusion
- Hugging Face Diffusers documentation: https://huggingface.co/docs/diffusers/index
About this lesson. This is the illustrated edition of a lesson from the open-source AI Primer. Its text, figures and numbers are generated from the Primer's source at commit c8d5c21, so the two always agree: the explanation, the code that builds it and the tests that prove it.