Visual explainer
Diffusion Models
Learn how diffusion models generate highly detailed images by iteratively reversing a noise-adding process.
Generating a highly detailed, photorealistic image directly from an empty canvas is an incredibly hard optimization problem for a neural network. Instead of trying to guess every pixel's exact color in a single shot, diffusion models take a radically different approach: they start with pure random static noise and carefully carve a structured image out of it over time.
Forward Process
Before the model can generate anything new, it must first learn exactly how noise corrupts structure. During training, a real image is put through a mathematical forward diffusion process. Tiny, mathematically precise amounts of Gaussian noise are added step by step. Over hundreds of iterations, the original visual structure is completely destroyed, leaving behind nothing but pure, unrecoverable static.
Reverse Denoising
The core mechanism is a neural network, typically a U-Net architecture, that learns to run this destructive process in reverse. Given a slightly noisy intermediate state, it predicts the exact noise pattern that was added during the forward step. By subtracting this predicted noise, the image becomes slightly clearer, repeatedly inching closer to reality with every sequential loop.
Text Conditioning
If we only blindly subtracted noise, we would get a random, uncontrollable image every time. To steer the output, text prompts are converted into rich embeddings and fed directly into the U-Net at every single step. This text conditioning powerfully guides the denoising process, forcing the chaotic noise to crystallize into the specific concepts requested by the user.
Where It Breaks
Because the reverse diffusion process relies on Markov chains—where each state depends entirely on the one immediately prior—it must happen sequentially. This often requires hundreds or thousands of iterative steps to produce a sharp, cohesive result. It is fundamentally slower than a single-pass generative network like a GAN, making real-time inference latency a very difficult hard limit.
The Quick Version
- The problem: Creating detailed images in one shot is too complex.
- Forward pass: Real images are systematically destroyed with Gaussian noise.
- Reverse pass: A U-Net learns to subtract that noise step by step.
- Conditioning: Text embeddings inject intent directly into the denoising loop.
- The catch: Sequential step-by-step generation inherently causes high inference latency.