Skip to content
AI360Xpert
Beta

Diffusion and Score-Based Models

Data is gradually corrupted with Gaussian noise over hundreds of steps, and a neural network learns to reverse the process step by step.

Diffusion models corrupt data with scheduled Gaussian noise, training a neural network to estimate the Stein score and denoise reverse trajectories.
Diffusion models corrupt data with scheduled Gaussian noise, training a neural network to estimate the Stein score and denoise reverse trajectories.

Why Does This Exist?

Prior to diffusion models, generative modeling was split between three flawed paradigms:

  1. Generative Adversarial Networks (GANs): Capable of generating sharp images, but notoriously unstable to train due to adversarial min-max dynamics, hyperparameter sensitivity, and catastrophic mode collapse.
  2. Variational Autoencoders (VAEs): Stable to train via likelihood bounds, but produce blurry images because approximating true posterior distributions with simple Gaussians forces the decoder to average across plausible modes.
  3. Autoregressive Models: Excellent likelihood estimates, but generate images one pixel or token at a time, resulting in severe inference latency.

Diffusion Models (Denoising Diffusion Probabilistic Models, DDPM) and Score-Based Generative Models (SGMs) overcome these limitations by framing generation as reversing a thermodynamic diffusion process. Rather than attempting to map a simple Gaussian distribution directly to complex natural images in a single massive jump, diffusion breaks generation down into hundreds or thousands of tiny, tractable denoising steps.

Training requires no adversarial discriminator; instead, a standard neural network (typically a U-Net) is trained with simple mean squared error to predict the noise injected at an arbitrary timestep. Training is unconditionally stable, covers all modes of the distribution, and achieves state-of-the-art sample quality. Foundation concepts rely on understanding regression loss-functions.

Think of It Like This

A perfume drop dissolving in water filmed in reverse

Imagine releasing a single drop of concentrated purple perfume into a clear glass of water.

In the forward process (diffusion), the concentrated drop naturally expands. Water molecules collide with perfume particles, scattering them randomly. At step 10, the drop is slightly fuzzy. At step 200, purple wisps curl through the liquid. At step 1,000, the glass is completely uniform, faint purple water: maximum entropy, all original structure completely destroyed. Every drop of perfume, regardless of whether it was shaped like a star or a circle, ends up as identical uniform noise.

Now imagine running the film backward. If you know the precise statistical laws governing how water molecules pushed the perfume particles outward, you can nudge every particle slightly inward against the current. By taking 1,000 microscopic reverse steps, you can start from completely random purple water and assemble the scattered molecules back into a pristine, concentrated drop.

How It Actually Works

Forward SDEs, Score Matching, and the DDPM Objective

The diffusion framework can be understood through two unified lenses: discrete Markov chains (DDPM) and continuous stochastic differential equations (Score-based models).

1. Discrete Forward Process (DDPM)

Given data x0∼q(x0)x_0 \sim q(x_0), the forward Markov chain adds Gaussian noise according to a fixed variance schedule β1,β2,…,βT\beta_1, \beta_2, \dots, \beta_T:

q(xt∣xt−1)=N(xt; 1−βtxt−1, βtI)q(x_t \mid x_{t-1}) = \mathcal{N}\left(x_t;\, \sqrt{1 - \beta_t} x_{t-1},\, \beta_t I\right)

A key mathematical property of Gaussian transitions is that we can sample xtx_t at any arbitrary timestep tt in closed form without iterating through intermediate steps. Defining αt=1−βt\alpha_t = 1 - \beta_t and αˉt=∏s=1tαs\bar{\alpha}_t = \prod_{s=1}^t \alpha_s:

q(xt∣x0)=N(xt; αˉtx0, (1−αˉt)I)q(x_t \mid x_0) = \mathcal{N}\left(x_t;\, \sqrt{\bar{\alpha}_t} x_0,\, (1 - \bar{\alpha}_t) I\right)

xt=αˉtx0+1−αˉtϵ,ϵ∼N(0,I)x_t = \sqrt{\bar{\alpha}_t} x_0 + \sqrt{1 - \bar{\alpha}_t} \epsilon, \quad \epsilon \sim \mathcal{N}(0, I)

2. Reverse Process and Simplified Loss

The reverse process pθ(xt−1∣xt)p_\theta(x_{t-1} \mid x_t) is also Gaussian for sufficiently small βt\beta_t:

pθ(xt−1∣xt)=N(xt−1; μθ(xt,t), Σθ(xt,t))p_\theta(x_{t-1} \mid x_t) = \mathcal{N}\left(x_{t-1};\, \mu_\theta(x_t, t),\, \Sigma_\theta(x_t, t)\right)

Ho et al. (2020) proved that reparameterizing the mean μθ\mu_\theta to predict the injected noise vector ϵ\epsilon yields a simplified variational lower bound objective:

Lsimple(θ)=Et∼U[1,T], x0∼q(x0), ϵ∼N(0,I)[∥ϵ−ϵθ(xt,t)∥22]\mathcal{L}_{\text{simple}}(\theta) = \mathbb{E}_{t \sim U[1, T],\, x_0 \sim q(x_0),\, \epsilon \sim \mathcal{N}(0, I)} \left[ \|\epsilon - \epsilon_\theta(x_t, t)\|_2^2 \right]

During training, an algorithm:

  1. Samples a clean image x0x_0 and a random timestep t∈{1,…,T}t \in \{1, \dots, T\}.
  2. Samples standard normal noise ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I).
  3. Computes noisy image xt=αˉtx0+1−αˉtϵx_t = \sqrt{\bar{\alpha}_t} x_0 + \sqrt{1 - \bar{\alpha}_t} \epsilon.
  4. Passes xtx_t and timestep embedding tt to neural network ϵθ\epsilon_\theta.
  5. Calculates MSE between true noise ϵ\epsilon and predicted noise ϵθ(xt,t)\epsilon_\theta(x_t, t).

3. Score-Based Perspective

Song and Ermon (2019) developed generation via the Stein score function, defined as the gradient of the log probability density with respect to the data:

s(x)≡∇xlog⁡p(x)s(x) \equiv \nabla_x \log p(x)

The score function is a vector field pointing in the direction where data density increases most rapidly. Score matching trains a neural network sθ(x,σ)s_\theta(x, \sigma) to approximate this gradient without ever calculating the intractable normalizing partition function of p(x)p(x).

The connection between DDPM and Score Matching is exact:

sθ(xt,t)=−ϵθ(xt,t)1−αˉts_\theta(x_t, t) = -\frac{\epsilon_\theta(x_t, t)}{\sqrt{1 - \bar{\alpha}_t}}

Sampling is performed using Annealed Langevin Dynamics:

xk+1=xk+δ2sθ(xk,σi)+δzk,zk∼N(0,I)x_{k+1} = x_k + \frac{\delta}{2} s_\theta(x_k, \sigma_i) + \sqrt{\delta} z_k, \quad z_k \sim \mathcal{N}(0, I)

Worked Example

Let us trace a single forward corruption and training evaluation step on a 2D toy vector x0=[1.2,−0.8]x_0 = [1.2, -0.8] at timestep t=500t = 500:

  1. Schedule parameter values at t=500t = 500: Suppose the cumulative noise factor is αˉ500=0.64\bar{\alpha}_{500} = 0.64. αˉ500=0.64=0.80\sqrt{\bar{\alpha}_{500}} = \sqrt{0.64} = 0.80 1−αˉ500=1−0.64=0.36=0.60\sqrt{1 - \bar{\alpha}_{500}} = \sqrt{1 - 0.64} = \sqrt{0.36} = 0.60

  2. Sample random noise ϵ\epsilon: Sample ϵ=[0.50,−1.00]∼N(0,I)\epsilon = [0.50, -1.00] \sim \mathcal{N}(0, I).

  3. Compute noisy coordinate x500x_{500}: x500=0.80×[1.2,−0.8]+0.60×[0.50,−1.00]x_{500} = 0.80 \times [1.2, -0.8] + 0.60 \times [0.50, -1.00] =[0.96,−0.64]+[0.30,−0.60]=[1.26,−1.24]= [0.96, -0.64] + [0.30, -0.60] = [1.26, -1.24]

  4. Model noise prediction and loss: Assume the U-Net evaluates x500x_{500} and outputs ϵ^=ϵθ(x500,500)=[0.45,−0.90]\hat{\epsilon} = \epsilon_\theta(x_{500}, 500) = [0.45, -0.90]. Compute squared error against true noise ϵ=[0.50,−1.00]\epsilon = [0.50, -1.00]: L=12((0.50−0.45)2+(−1.00−(−0.90))2)\mathcal{L} = \frac{1}{2}\left( (0.50 - 0.45)^2 + (-1.00 - (-0.90))^2 \right) =12(0.052+(−0.10)2)=12(0.0025+0.0100)=0.00625= \frac{1}{2}\left( 0.05^2 + (-0.10)^2 \right) = \frac{1}{2}(0.0025 + 0.0100) = 0.00625

  5. Corresponding score vector estimate: sθ(x500,500)=−[0.45,−0.90]0.60=[−0.75,1.50]s_\theta(x_{500}, 500) = -\frac{[0.45, -0.90]}{0.60} = [-0.75, 1.50] The score vector points back toward the high-density region where x0=[1.2,−0.8]x_0 = [1.2, -0.8] resides!

Code

import torchimport torch.nn as nnimport torch.nn.functional as F
class DiffusionForward:    """Precomputes closed-form coefficients for forward Gaussian diffusion."""    def __init__(self, timesteps: int = 1000, beta_start: float = 1e-4, beta_end: float = 0.02):        self.timesteps = timesteps        # Linear beta schedule        self.betas = torch.linspace(beta_start, beta_end, timesteps)        self.alphas = 1.0 - self.betas        self.alphas_cumprod = torch.cumprod(self.alphas, dim=0)        self.sqrt_alphas_cumprod = torch.sqrt(self.alphas_cumprod)        self.sqrt_one_minus_alphas_cumprod = torch.sqrt(1.0 - self.alphas_cumprod)
    def q_sample(self, x_0: torch.Tensor, t: torch.Tensor, noise: torch.Tensor | None = None) -> tuple[torch.Tensor, torch.Tensor]:        """Closed-form jump to arbitrary timestep t."""        if noise is None:            noise = torch.randn_like(x_0)                sqrt_alpha = self.sqrt_alphas_cumprod[t].view(-1, 1, 1, 1)        sqrt_one_minus_alpha = self.sqrt_one_minus_alphas_cumprod[t].view(-1, 1, 1, 1)                x_t = sqrt_alpha * x_0 + sqrt_one_minus_alpha * noise        return x_t, noise
# Test closed-form forward samplingdiff = DiffusionForward(timesteps=1000)x_0 = torch.randn(4, 3, 32, 32)# Sample random timesteps for a batcht = torch.tensor([50, 250, 500, 900])x_t, noise = diff.q_sample(x_0, t)
print("x_t shape:", x_t.shape)# -> x_t shape: torch.Size([4, 3, 32, 32])print("Noise shape:", noise.shape)# -> Noise shape: torch.Size([4, 3, 32, 32])
# Measure variance at t=50 vs t=900alpha_bar_50 = diff.alphas_cumprod[50].item()alpha_bar_900 = diff.alphas_cumprod[900].item()print(f"Signal preserved at t=50: {alpha_bar_50 * 100:.1f}%")# -> Signal preserved at t=50: 98.4%print(f"Signal preserved at t=900: {alpha_bar_900 * 100:.1f}%")# -> Signal preserved at t=900: 0.8%

Watch Out For

Sampling timesteps uniformly without noise schedule variance adjustment

When training on higher-resolution images or using linear beta schedules (β1=10−4\beta_1 = 10^{-4} to βT=0.02\beta_T = 0.02), the signal-to-noise ratio drops to near-zero too quickly in the middle of the trajectory. By timestep t=500t=500, virtually all high-frequency details are gone, forcing the network to spend over half its training capacity predicting pure noise without learning structural semantic features.

In modern implementations, switch from a naive linear schedule to a cosine noise schedule (Nichol & Dhariwal, 2021). The cosine schedule ensures that αˉt\bar{\alpha}_t drops smoothly following cos⁡2(t/T+s1+s⋅π2)\cos^2\left(\frac{t/T + s}{1 + s} \cdot \frac{\pi}{2}\right), preserving useful image structural information across a much wider fraction of timesteps and dramatically accelerating convergence.

The Quick Version

  • Diffusion models decompose generative modeling into hundreds of small denoising steps, eliminating the training instability and mode collapse inherent to GANs.
  • The forward process adds Gaussian noise according to variance schedule βt\beta_t, permitting direct closed-form sampling of noisy state xtx_t at any timestep tt.
  • Training minimizes simple MSE loss between true added noise ϵ\epsilon and network prediction ϵθ(xt,t)\epsilon_\theta(x_t, t), which is mathematically equivalent to estimating the Stein score ∇xlog⁡p(x)\nabla_x \log p(x).