Visual explainer
Vanishing & Exploding Gradients
Why deep networks fail to learn. See how backpropagation multiplies errors, shrinking them to zero or blowing them up to infinity, and how modern networks fix it.
To update weights in early layers, the error signal must propagate backward through every subsequent layer. Because of the chain rule, this means multiplying the error by the derivative of each layer. If the activation function (like Sigmoid) has a maximum derivative of 0.25, the gradient is multiplied by a fraction at every step, shrinking exponentially until the early layers freeze completely.
Exploding Gradients
The reverse problem happens when the network's weights or derivatives are consistently greater than 1. Instead of decaying, the error signal amplifies exponentially backward through the network. This causes the gradients to explode, leading to chaotic, massive weight updates that destabilize training, often resulting in "NaN" loss values.
The Modern Fixes
Modern architectures rely on specific structural fixes rather than just careful tuning. To prevent vanishing gradients, Residual Connections (ResNets) add a direct "superhighway" that passes the gradient straight through, bypassing the fraction-multiplying layers. To prevent exploding gradients, Gradient Clipping imposes a hard mathematical ceiling on the magnitude of the gradient, keeping it safe and bounded.
The Quick Version
- The chain rule requires multiplying derivatives to pass errors backward.
- Vanishing: Successive multiplication by small numbers (e.g., Sigmoid's 0.25) shrinks the gradient, freezing early layers.
- Exploding: Successive multiplication by large numbers amplifies the gradient, causing chaotic updates or overflow.
- Solutions: Residual connections and ReLU prevent vanishing, while gradient clipping prevents exploding.