Learning Rate Warmup
Increasing the learning rate from near zero up to its target value over the first steps of training, before any decay schedule takes over.
Adam's second-moment estimate starts at zero and needs a handful of steps of bias correction before it's a trustworthy read on a parameter's gradient scale. Applying a full-strength learning rate against that unstable early estimate is the specific combination that can spike the loss or send it to nan in the first few hundred steps.
Warmup ramps the rate up linearly instead, bounding how much damage those early, unreliable steps can do. It's close to standard practice for training transformers from scratch, and less critical for smaller models trained with plain SGD.