Denoising Diffusion Implicit Models (DDIM)
By generalizing diffusion to a non-Markovian forward process, generation becomes a deterministic trajectory that cuts sampling steps by twenty-fold.
Why Does This Exist?
Standard Denoising Diffusion Probabilistic Models (DDPM) unlocked unprecedented visual fidelity, but they suffered from an crippling operational bottleneck: inference latency. Because DDPM defines its reverse generation process as a discrete Markov chain, generating a single image requires sequentially evaluating the denoising U-Net for steps. On a high-end GPU, generating one image took between 10 and 30 seconds.
Furthermore, because DDPM injects independent Gaussian noise at every single reverse step (), the trajectory is strictly stochastic. If a user feeds an existing photograph into the model, there is no deterministic path to invert the photograph back into a unique latent noise representation. Without an invertible latent space, tasks like deterministic image editing, semantic feature blending, and latent-space interpolation become impossible.
Denoising Diffusion Implicit Models (DDIM) (Song et al., 2020) resolved both problems with a profound theoretical insight: the forward noising process does not need to be Markovian to share the exact same marginal distributions as DDPM. By designing a non-Markovian forward process parameterized by a stochasticity coefficient , setting converts the generation trajectory into a deterministic Ordinary Differential Equation (ODE).
This deterministic formulation allows inference to skip 95% of steps—generating high-fidelity images in just 20 to 50 steps using the exact same pretrained DDPM weights with zero retraining—while establishing a bidirectional bijection between images and noise vectors. Conceptual continuity builds directly upon diffusion-score-based-models.
Think of It Like This
A mountain descent: drifting through a fog versus driving along surveyed contour lines
Imagine descending from the summit of a mountain (, pure noise) to a specific village in the valley (, a clean image).
DDPM is like a hiker walking downhill through thick fog. At every single meter of the descent, the hiker checks a compass for general direction, but then stumbles and takes a random sideways step driven by mountain wind (stochastic noise ). Because there are 1,000 random gusts of wind along the descent, the hiker must take tiny, cautious steps to avoid falling off a cliff. If the hiker reaches a village and tries to climb back up to the exact same summit spot, the random wind makes tracing the path in reverse completely impossible.
DDIM is like a highway engineer paving a paved, banked road down the mountain along the smooth contour lines of the terrain (). Because the road follows a continuous, smooth differential curve, you don't need to stop every meter: you can set your cruise control and check your navigation at only 20 milestones along the descent. Even better, because the road is paved and deterministic, you can turn your car around at the valley, drive up the exact same highway in reverse, and arrive at the exact unique parking spot on the summit.
How It Actually Works
Non-Markovian Forward Processes and the Deterministic Reverse Step
In standard DDPM, the forward distribution is strictly Markovian: . DDIM generalizes this by considering a family of non-Markovian forward distributions that satisfy two critical conditions:
- The marginal distribution at every step remains identical to DDPM:
- The joint distribution conditions on both the previous state and the clean origin :
Because the marginals are identical to DDPM, the objective function is also identical. Any neural network trained for DDPM can be evaluated with DDIM sampling with zero modification.
The general reverse step equation from timestep to (or along an arbitrary sub-sequence to ) is:
where the stochasticity hyperparameter is defined as:
Two Critical Regimes of :
- When : equals the forward posterior variance of DDPM, recovering standard stochastic Markovian diffusion.
- When : . The random noise term vanishes completely!
When , the reverse update becomes a deterministic discretization of the Probability Flow ODE:
Sub-sequence Striding
Because the transition is governed by an ODE rather than a step-by-step Markov chain, we can select an arbitrary strided sub-sequence of timesteps where . For example, choosing steps out of evaluates timesteps , cutting inference computation by with negligible loss in sample quality.
Worked Example
Let us trace a single deterministic DDIM step () jumping from timestep down to :
-
Given schedule parameters:
- At step : , .
- At step : , .
-
Current state and model prediction: Suppose current scalar coordinate is , and the neural network predicts noise .
-
Step 1: Estimate predicted clean origin :
-
Step 2: Compute next deterministic state :
Notice: Zero random numbers were sampled. The mapping from to is completely deterministic and exact.
Code
import torch
def ddim_step( x_t: torch.Tensor, eps_pred: torch.Tensor, alpha_bar_t: float, alpha_bar_prev: float, eta: float = 0.0,) -> torch.Tensor: """Computes a single reverse DDIM step from timestep t to prev_t. When eta=0.0, the update is completely deterministic. """ # 1. Predict clean x_0 from current noisy state sqrt_alpha_bar_t = alpha_bar_t ** 0.5 sqrt_one_minus_alpha_bar_t = (1.0 - alpha_bar_t) ** 0.5 pred_x0 = (x_t - sqrt_one_minus_alpha_bar_t * eps_pred) / sqrt_alpha_bar_t # 2. Compute stochastic variance sigma_t if eta > 0.0: sigma = eta * ( ((1.0 - alpha_bar_prev) / (1.0 - alpha_bar_t)) ** 0.5 * (1.0 - alpha_bar_t / alpha_bar_prev) ** 0.5 ) else: sigma = 0.0 # 3. Direction vector pointing to x_t dir_xt = (1.0 - alpha_bar_prev - sigma ** 2) ** 0.5 * eps_pred # 4. Deterministic next state x_prev = (alpha_bar_prev ** 0.5) * pred_x0 + dir_xt # 5. Add stochastic noise if eta > 0 if sigma > 0.0: noise = torch.randn_like(x_t) x_prev = x_prev + sigma * noise return x_prev
# Demonstration: Deterministic reproducibilityx_current = torch.tensor([1.50])eps_model = torch.tensor([0.60])a_bar_t = 0.70a_bar_prev = 0.85
# Running the function twice with eta=0.0 yields identical tensorsout1 = ddim_step(x_current, eps_model, a_bar_t, a_bar_prev, eta=0.0)out2 = ddim_step(x_current, eps_model, a_bar_t, a_bar_prev, eta=0.0)
print(f"DDIM output: {out1.item():.4f}")# -> DDIM output: 1.5231print("Exact equality across independent runs:", torch.equal(out1, out2))# -> Exact equality across independent runs: TrueWatch Out For
Accumulated discretization drift during DDIM inversion without null-text guidance matching
Inverting a real image to its latent noise vector via reverse ODE integration relies on local Euler approximations: . When running reverse ODE generation from back to , discretization errors compound at each step. If Classifier-Free Guidance (CFG) is enabled during generation, the forward inversion and reverse generation trajectories diverge significantly, resulting in generated images that alter facial identities or structural compositions.
To achieve exact reconstruction in image editing pipelines, perform Null-Text Inversion (Mokady et al., 2023): keep the model guidance scale aligned or optimize the unconditional text embedding vector at each inversion step to guarantee that the reverse ODE trajectory retraces the forward inversion trajectory with near-zero mathematical drift.
The Quick Version
- DDIM reformulates the diffusion forward process as non-Markovian, preserving the exact same marginal distributions and training objective as DDPM.
- Setting the stochasticity parameter transforms diffusion into a deterministic Probability Flow ODE.
- Strided sub-sequence sampling allows generating high-fidelity images in 20 to 50 steps ( faster than DDPM) while enabling exact latent space inversion for image editing.