Skip to content
AI360Xpert
Core ML

Visual explainer

Vanishing & Exploding Gradients

Why deep networks fail to learn. See how backpropagation multiplies errors, shrinking them to zero or blowing them up to infinity, and how modern networks fix it.

The chain rule multiplies gradients layer by layer, shrinking the signal to zero.
The chain rule multiplies gradients layer by layer, shrinking the signal to zero.

To update weights in early layers, the error signal must propagate backward through every subsequent layer. Because of the chain rule, this means multiplying the error by the derivative of each layer. If the activation function (like Sigmoid) has a maximum derivative of 0.25, the gradient is multiplied by a fraction at every step, shrinking exponentially until the early layers freeze completely.

Exploding Gradients

Unconstrained large weights can amplify the gradient exponentially, destroying the model.
Unconstrained large weights can amplify the gradient exponentially, destroying the model.

The reverse problem happens when the network's weights or derivatives are consistently greater than 1. Instead of decaying, the error signal amplifies exponentially backward through the network. This causes the gradients to explode, leading to chaotic, massive weight updates that destabilize training, often resulting in "NaN" loss values.

The Modern Fixes

Residual connections bypass layers for healthy gradients, and clipping caps explosive growth.
Residual connections bypass layers for healthy gradients, and clipping caps explosive growth.

Modern architectures rely on specific structural fixes rather than just careful tuning. To prevent vanishing gradients, Residual Connections (ResNets) add a direct "superhighway" that passes the gradient straight through, bypassing the fraction-multiplying layers. To prevent exploding gradients, Gradient Clipping imposes a hard mathematical ceiling on the magnitude of the gradient, keeping it safe and bounded.

The Quick Version

  • The chain rule requires multiplying derivatives to pass errors backward.
  • Vanishing: Successive multiplication by small numbers (e.g., Sigmoid's 0.25) shrinks the gradient, freezing early layers.
  • Exploding: Successive multiplication by large numbers amplifies the gradient, causing chaotic updates or overflow.
  • Solutions: Residual connections and ReLU prevent vanishing, while gradient clipping prevents exploding.