Diffusion and Score-Based Models
Data is gradually corrupted with Gaussian noise over hundreds of steps, and a neural network learns to reverse the process step by step.
Why Does This Exist?
Prior to diffusion models, generative modeling was split between three flawed paradigms:
- Generative Adversarial Networks (GANs): Capable of generating sharp images, but notoriously unstable to train due to adversarial min-max dynamics, hyperparameter sensitivity, and catastrophic mode collapse.
- Variational Autoencoders (VAEs): Stable to train via likelihood bounds, but produce blurry images because approximating true posterior distributions with simple Gaussians forces the decoder to average across plausible modes.
- Autoregressive Models: Excellent likelihood estimates, but generate images one pixel or token at a time, resulting in severe inference latency.
Diffusion Models (Denoising Diffusion Probabilistic Models, DDPM) and Score-Based Generative Models (SGMs) overcome these limitations by framing generation as reversing a thermodynamic diffusion process. Rather than attempting to map a simple Gaussian distribution directly to complex natural images in a single massive jump, diffusion breaks generation down into hundreds or thousands of tiny, tractable denoising steps.
Training requires no adversarial discriminator; instead, a standard neural network (typically a U-Net) is trained with simple mean squared error to predict the noise injected at an arbitrary timestep. Training is unconditionally stable, covers all modes of the distribution, and achieves state-of-the-art sample quality. Foundation concepts rely on understanding regression loss-functions.
Think of It Like This
A perfume drop dissolving in water filmed in reverse
Imagine releasing a single drop of concentrated purple perfume into a clear glass of water.
In the forward process (diffusion), the concentrated drop naturally expands. Water molecules collide with perfume particles, scattering them randomly. At step 10, the drop is slightly fuzzy. At step 200, purple wisps curl through the liquid. At step 1,000, the glass is completely uniform, faint purple water: maximum entropy, all original structure completely destroyed. Every drop of perfume, regardless of whether it was shaped like a star or a circle, ends up as identical uniform noise.
Now imagine running the film backward. If you know the precise statistical laws governing how water molecules pushed the perfume particles outward, you can nudge every particle slightly inward against the current. By taking 1,000 microscopic reverse steps, you can start from completely random purple water and assemble the scattered molecules back into a pristine, concentrated drop.
How It Actually Works
Forward SDEs, Score Matching, and the DDPM Objective
The diffusion framework can be understood through two unified lenses: discrete Markov chains (DDPM) and continuous stochastic differential equations (Score-based models).
1. Discrete Forward Process (DDPM)
Given data , the forward Markov chain adds Gaussian noise according to a fixed variance schedule :
A key mathematical property of Gaussian transitions is that we can sample at any arbitrary timestep in closed form without iterating through intermediate steps. Defining and :
2. Reverse Process and Simplified Loss
The reverse process is also Gaussian for sufficiently small :
Ho et al. (2020) proved that reparameterizing the mean to predict the injected noise vector yields a simplified variational lower bound objective:
During training, an algorithm:
- Samples a clean image and a random timestep .
- Samples standard normal noise .
- Computes noisy image .
- Passes and timestep embedding to neural network .
- Calculates MSE between true noise and predicted noise .
3. Score-Based Perspective
Song and Ermon (2019) developed generation via the Stein score function, defined as the gradient of the log probability density with respect to the data:
The score function is a vector field pointing in the direction where data density increases most rapidly. Score matching trains a neural network to approximate this gradient without ever calculating the intractable normalizing partition function of .
The connection between DDPM and Score Matching is exact:
Sampling is performed using Annealed Langevin Dynamics:
Worked Example
Let us trace a single forward corruption and training evaluation step on a 2D toy vector at timestep :
-
Schedule parameter values at : Suppose the cumulative noise factor is .
-
Sample random noise : Sample .
-
Compute noisy coordinate :
-
Model noise prediction and loss: Assume the U-Net evaluates and outputs . Compute squared error against true noise :
-
Corresponding score vector estimate: The score vector points back toward the high-density region where resides!
Code
import torchimport torch.nn as nnimport torch.nn.functional as F
class DiffusionForward: """Precomputes closed-form coefficients for forward Gaussian diffusion.""" def __init__(self, timesteps: int = 1000, beta_start: float = 1e-4, beta_end: float = 0.02): self.timesteps = timesteps # Linear beta schedule self.betas = torch.linspace(beta_start, beta_end, timesteps) self.alphas = 1.0 - self.betas self.alphas_cumprod = torch.cumprod(self.alphas, dim=0) self.sqrt_alphas_cumprod = torch.sqrt(self.alphas_cumprod) self.sqrt_one_minus_alphas_cumprod = torch.sqrt(1.0 - self.alphas_cumprod)
def q_sample(self, x_0: torch.Tensor, t: torch.Tensor, noise: torch.Tensor | None = None) -> tuple[torch.Tensor, torch.Tensor]: """Closed-form jump to arbitrary timestep t.""" if noise is None: noise = torch.randn_like(x_0) sqrt_alpha = self.sqrt_alphas_cumprod[t].view(-1, 1, 1, 1) sqrt_one_minus_alpha = self.sqrt_one_minus_alphas_cumprod[t].view(-1, 1, 1, 1) x_t = sqrt_alpha * x_0 + sqrt_one_minus_alpha * noise return x_t, noise
# Test closed-form forward samplingdiff = DiffusionForward(timesteps=1000)x_0 = torch.randn(4, 3, 32, 32)# Sample random timesteps for a batcht = torch.tensor([50, 250, 500, 900])x_t, noise = diff.q_sample(x_0, t)
print("x_t shape:", x_t.shape)# -> x_t shape: torch.Size([4, 3, 32, 32])print("Noise shape:", noise.shape)# -> Noise shape: torch.Size([4, 3, 32, 32])
# Measure variance at t=50 vs t=900alpha_bar_50 = diff.alphas_cumprod[50].item()alpha_bar_900 = diff.alphas_cumprod[900].item()print(f"Signal preserved at t=50: {alpha_bar_50 * 100:.1f}%")# -> Signal preserved at t=50: 98.4%print(f"Signal preserved at t=900: {alpha_bar_900 * 100:.1f}%")# -> Signal preserved at t=900: 0.8%Watch Out For
Sampling timesteps uniformly without noise schedule variance adjustment
When training on higher-resolution images or using linear beta schedules ( to ), the signal-to-noise ratio drops to near-zero too quickly in the middle of the trajectory. By timestep , virtually all high-frequency details are gone, forcing the network to spend over half its training capacity predicting pure noise without learning structural semantic features.
In modern implementations, switch from a naive linear schedule to a cosine noise schedule (Nichol & Dhariwal, 2021). The cosine schedule ensures that drops smoothly following , preserving useful image structural information across a much wider fraction of timesteps and dramatically accelerating convergence.
The Quick Version
- Diffusion models decompose generative modeling into hundreds of small denoising steps, eliminating the training instability and mode collapse inherent to GANs.
- The forward process adds Gaussian noise according to variance schedule , permitting direct closed-form sampling of noisy state at any timestep .
- Training minimizes simple MSE loss between true added noise and network prediction , which is mathematically equivalent to estimating the Stein score .