Hard Negative Mining Training
Hard-negative mining spends training batches on the confusing near-misses instead of the easy cases the model already gets right.
Why Does This Exist?
Random batches are mostly trivia: strangers already far apart, backgrounds obviously empty. Losses on easy samples sit at zero and gradients vanish while the GPU burns. Training stalls with the decision boundary still sloppy exactly where it matters. Hard-negative mining exists to aim each batch at the current confusion frontier: the negatives sitting closest to anchors, the background boxes scoring highest. Every page on contrastive loss, triplet loss and detection depends on it.
Think of It Like This
A coach who replays your errors
A tennis coach who feeds you only easy lobs wastes the lesson; one who replays the backhands you dumped into the net fixes your game. Mining is the coach's tape review: rank recent mistakes by embarrassment and drill those. The analogy stops at the danger both share: drilling only the single worst error teaches twitchy overcorrection, which is why semi-hard selection beats hardest-only.
How It Actually Works
Online mining scores all candidate negatives in the batch with the current model and keeps the top offenders: highest-loss background boxes for detectors (OHEM), closest-to-anchor negatives for triplets. Semi-hard variants keep negatives inside the margin but still farther than the positive, which trains steadily on noisy labels. Offline mining refreshes a cached hard set every few epochs for retrieval at scale.
A worked triplet pick
An anchor's batch offers negatives at distances 1.4, 0.45 and 0.32 with the positive at 0.30 and margin 0.2. The 1.4 negative is easy (loss 0) and teaches nothing. The 0.32 negative is hard (loss 0.18) but possibly mislabelled. The 0.45 negative is semi-hard: loss max(0, 0.30 - 0.45 + 0.2) = 0.05, a safe steady gradient. Semi-hard mining trains on the 0.45 case and skips both extremes.
Watch Out For
Hardest-only mining that memorizes noise
The single hardest negatives are disproportionately mislabelled, and training on them drags the boundary toward errors. The symptom is training loss that thrashes while validation retrieval decays. Fix it with semi-hard ranges, label cleaning, and capping the hardest fraction per batch.
The Quick Version
- Easy samples give zero gradients, so random batches waste training.
- Mining keeps the highest-loss negatives per batch for detectors and embeddings.
- Semi-hard selection trains steadily where hardest-only chases label noise.
- Refresh mined sets as the model improves; yesterday's hards are today's easies.
- Budget mining ratios explicitly or batches collapse into pure noise drills.