Stochastic Gradient Descent Variants
Stochastic gradient descent estimates the true dataset gradient from tiny random mini-batches, using momentum and adaptive coordinate scaling to accelerate training.
Why Does This Exist?
In classical batch gradient descent, computing a single parameter update requires evaluating the loss gradient across every single training example in the dataset. If your dataset contains images, taking one optimization step requires ten million forward passes and ten million backward passes through the neural network. Training would grind to an absolute halt.
Stochastic Gradient Descent (SGD) solves this bottleneck through random sampling: instead of computing the true full-batch gradient over all examples, it computes an unbiased gradient estimate over a small mini-batch of samples (typically to ). This accelerates iteration throughput by factors of thousands.
However, mini-batch estimation introduces substantial gradient variance (noise) and struggles with ill-conditioned loss landscapes. In narrow loss ravines where curvature is steep along one parameter axis and shallow along another, vanilla SGD ricochets violently between the steep ravine walls while making negligible progress along the gentle valley floor.
To overcome this, an evolutionary family of optimizer variants emerged: Momentum adds physical inertia to cancel out oscillating noise, RMSprop normalizes updates by historical gradient magnitudes, and Adam synthesizes momentum and adaptive variance scaling into the modern default optimizer for deep learning.
For the fundamentals of directional derivatives and loss slopes, see our guide on gradients.
Think of It Like This
A ping-pong ball vs. a heavy bowling ball with studded snow treads
Imagine trying to roll a ball down a steep, icy half-pipe ravine to reach the exit at the bottom. The sides of the half-pipe rise sharply to your left and right, but the ravine floor tilts only gently forward toward the finish line.
Vanilla SGD is like a lightweight ping-pong ball buffeted by erratic wind gusts (mini-batch noise). When you release it, gravity yanks it down the steep side wall. Because it has no mass or memory, it shoots straight across the floor and ricochets wildly up the opposite wall. It bounces endlessly back and forth across the ravine, traveling miles laterally while advancing only inches forward down the floor.
SGD with Momentum turns the ping-pong ball into a heavy bowling ball. Lateral bounces cancel each other out across steps, while consistent forward gravitational pull accumulates velocity, barreling smoothly down the center of the half-pipe.
Adam equips that bowling ball with intelligent motorized treads on each coordinate: it automatically applies brakes along the steep bouncing axis while hitting the gas pedal along the gentle, quiet forward axis.
The analogy stops because physical bowling balls obey continuous Newtonian physics in 3 dimensions with constant gravitational acceleration, whereas neural optimizers operate in discrete steps across non-Euclidean parameter landscapes.
How It Actually Works
From Mini-Batch SGD to Adaptive Moments
Given total training dataset and model parameter vector , the true loss is the empirical expectation:
1. Mini-Batch Stochastic Gradient Descent
At step , we draw a random mini-batch of size . The mini-batch gradient is:
The stochastic update is:
Because samples are chosen uniformly at random, is an unbiased estimator of the true gradient: . However, its covariance matrix scales inversely with batch size: .
2. Polyak Momentum
To suppress variance and accelerate through ravines, Polyak momentum introduces a velocity vector governed by friction coefficient :
Expanding this recurrence reveals an exponential moving average of past gradients:
Along oscillating axes, gradients flip signs (), summing toward zero. Along consistent directions, gradients reinforce each other, multiplying the effective step size by (a acceleration when ).
3. RMSprop (Root Mean Square Propagation)
In anisotropic landscapes, different parameters require drastically different learning rates. RMSprop maintains an exponential moving average of the squared gradients :
where denotes element-wise squaring. The update normalizes the step coordinate-wise:
where denotes element-wise multiplication and prevents division by zero. Parameters with large historical gradients are scaled down; parameters with tiny gradients are amplified.
4. Adam (Adaptive Moment Estimation)
Adam combines the first moment (momentum) and second raw moment (RMSprop) while introducing crucial initialization bias corrections:
Because and , both vectors are biased toward zero at early steps. Unbiasing yields:
The final parameter update is:
Standard hyperparameter defaults are , , and .
Worked Example
Let us compute the first step () of the Adam optimizer by hand.
Given parameters , learning rate , , , and . Suppose the observed gradient at is severely ill-conditioned:
The slope along is twenty times steeper than along .
-
First moment update ():
-
Second moment update ():
-
Bias correction at step :
-
Effective parameter update:
Notice the outcome: even though the raw gradient along was larger than along , Adam normalized both displacement components to , completely neutralizing the ill-conditioned ravine on step 1!
Code
import numpy as np
class AdamOptimizer:
def __init__( self, lr: float = 0.1, beta1: float = 0.9, beta2: float = 0.999, eps: float = 1e-8, ) -> None: self.lr = lr self.beta1 = beta1 self.beta2 = beta2 self.eps = eps self.m: np.ndarray | None = None self.v: np.ndarray | None = None self.t = 0
def step(self, theta: np.ndarray, grad: np.ndarray) -> np.ndarray: if self.m is None: self.m = np.zeros_like(theta) self.v = np.zeros_like(theta)
self.t += 1 # Update biased 1st and 2nd moment estimates self.m = self.beta1 * self.m + (1.0 - self.beta1) * grad self.v = self.beta2 * self.v + (1.0 - self.beta2) * (grad**2)
# Compute bias-corrected moments m_hat = self.m / (1.0 - self.beta1**self.t) v_hat = self.v / (1.0 - self.beta2**self.t)
# Apply adaptive update theta_next = theta - self.lr * m_hat / (np.sqrt(v_hat) + self.eps) return theta_next
# Test on ill-conditioned gradientopt = AdamOptimizer(lr=0.1)theta_0 = np.array([1.0, 2.0])grad_1 = np.array([0.2, 4.0])
theta_1 = opt.step(theta_0, grad_1)print(f"Step 1 theta: {theta_1}")# -> Step 1 theta: [0.9 1.9]
# Verify step 2 with alternating gradient signgrad_2 = np.array([0.2, -3.8])theta_2 = opt.step(theta_1, grad_2)print(f"Step 2 theta: {theta_2}")# -> Step 2 theta: [0.8 1.90481546]Watch Out For
Omitting Adam's bias correction during warm-up
Failing to implement Adam's bias corrections ( and ) is a severe implementation error.
Because is initialized at zero and , the second moment without correction at step is . Dividing by produces an effective step scale roughly times larger than intended! This uncalibrated spike destabilizes transformers and language models in their opening steps.
Fix: Always apply exact bias-corrected estimates and , and pair Adam with a learning rate warm-up schedule over the first 500 to 2,000 steps of training.
The Quick Version
- Mini-batch SGD computes unbiased gradient estimates over subsets of size , trading per-step precision for massive data throughput.
- Momentum accumulates past velocity vectors with decay , damping high-frequency lateral oscillations and accelerating through ravines.
- RMSprop divides updates by the root mean square of historical gradients, equalizing step sizes across parameters with mismatched scales.
- Adam combines first-moment momentum with second-moment variance scaling and exact bias correction, providing the standard robust optimizer for deep architectures.
- Always use learning rate warm-up schedules when training deep networks with Adam to prevent early gradient variance from destabilizing attention layers.