SphereFace Angular Softmax Loss
SphereFace rewrites softmax with angles only and multiplies the target angle, teaching faces to huddle tightly on a hypersphere.
Why Does This Exist?
Before margins, face networks trained plain softmax on raw dot products, mixing vector length with direction so a long mediocre vector could outscore a short excellent one. Verification on unseen people inherited the mess. SphereFace (A-Softmax) exists as the first fix: normalize away magnitude, decide purely by angle, and multiply the target angle by m so classes must separate angularly. It opened the path that CosFace and ArcFace later paved more smoothly.
Think of It Like This
Judging archers by angle, not arrow length
Judging archers by where long arrows land rewards strong arms over steady aim. Judging by launch angle alone rewards aim: identical angles score identically whatever the bow's power. SphereFace judges embeddings by angle alone, then demands the winner's angle be m times better than the runner-up's. The analogy stops at the difficulty: multiplicative margins make optimization twitchy, which is why later losses switched to gentler additive margins.
How It Actually Works
Weights are normalized so each logit is vector length times cosine of the angle to a class direction. The target logit uses cos(m x theta) instead of cos(theta) with integer m (often 4), implemented with Chebyshev polynomials to stay differentiable. Because cos(m x theta) drops far faster than cos(theta) as angles grow, the target class must sit dramatically closer than any rival to win.
A worked demand
With m = 4, a face 10 degrees from its class center scores cos(40), about 0.766, times its length. A rival 30 degrees away scores cos(30), about 0.866, and wins despite being three times farther angularly. The face must close to about 7 degrees (cos(28) = 0.883) to retake the lead. Small angular sloppiness costs large logit drops, which is the tightening pressure in one number.
Watch Out For
Training instability from the multiplicative margin
The cos(m x theta) surface oscillates for large angles, so early training with random weights can diverge where additive margins converge. The symptom is loss that explodes in the first epochs while CosFace trains fine. Fix it by annealing m from 1 upward or by warming up with plain normalized softmax first.
The Quick Version
- SphereFace decides by angle alone after normalizing class weights.
- Multiplying the target angle by m punishes angular sloppiness steeply.
- It pioneered angular margins; additive successors train more stably.
- Anneal m from 1 or warm up to avoid early divergence.
- Magnitude-free scoring is the idea that survived into every later margin loss.