Triplet Loss Anchor Training
Triplet loss trains on triples of anchor, match and stranger, demanding the match beat the stranger by a margin every single time.
Why Does This Exist?
Contrastive loss judges pairs in isolation with one global margin, which struggles when some identities naturally vary more than others. A relative rule works better: whatever the scale, the match must beat the stranger. Triplet loss exists as that relative rule, and it made FaceNet practical. Each training unit carries its own context, an anchor with both a friend and a stranger, so the gradient always knows what "closer" currently means.
Think of It Like This
A judge at a lookalike contest
A judge holds the anchor's photo, stands one contestant claiming to be the twin on the left and a stranger on the right, and demands the twin stand clearly closer. The judge never measures absolute distance to the stage: only the gap between the two contestants matters, and ties go to retraining. Triplets judge embeddings the same way. The analogy stops at the margin: the judge's "clearly" is a number added to the inequality, enforced by gradient descent.
How It Actually Works
For anchor a, positive p and negative n with distances d(a,p) and d(a,n), the loss is max(0, d(a,p) - d(a,n) + margin). Easy triplets where the stranger already trails by more than the margin cost zero. Semi-hard triplets where the stranger sits inside the margin but still farther than the match teach steadily; hard triplets where the stranger currently wins teach fastest but noisily.
A worked triplet
Use squared distances with margin 0.2. A semi-hard case has d(a,p) = 0.30 and d(a,n) = 0.40, giving loss max(0, 0.30 - 0.40 + 0.2) = 0.10: a live gradient. An easy case with d(a,p) = 0.05 and d(a,n) = 0.40 gives max(0, 0.05 - 0.40 + 0.2) = 0: nothing to learn. Online mining assembles each batch from the semi-hard and hard cases, because batches of easy triplets burn compute for zero gradient.
Watch Out For
Hard-only mining that chases label noise
The hardest triplets in a noisy dataset are mislabelled pairs, and training on them exclusively pulls the space toward the noise. The symptom is oscillating loss with degrading validation retrieval. Fix it with semi-hard mining: train on negatives inside the margin but still farther than the positive, and clean the noisiest labels first.
The Quick Version
- Triplet loss enforces a relative order per triplet, not an absolute distance.
- Easy triplets cost zero; semi-hard and hard triplets supply the gradient.
- Online mining inside each batch is mandatory because random triplets are trivial.
- Hard-only mining overfits label noise, so prefer semi-hard sampling on dirty data.
- The margin sets the required gap between match and stranger distances.