ArcFace Angular Margin Loss
ArcFace trains face embeddings on the surface of a sphere and demands an angular safety gap between identities, which sharpens open-set matching.
Why Does This Exist?
Plain softmax separates training identities but leaves their embedding clusters sprawling: the boundary sits wherever is convenient, and unseen people at test time land in the gaps. Open-set recognition needs compact clusters with enforced gaps, not just separable ones. ArcFace exists to carve those gaps during ordinary classification training, no triplet mining required. It keeps the softmax machinery and adds a geometric demand: each identity must win by an angular margin. This page covers the margin mechanism; the embedding lookup around it is in FaceNet and the pipeline in face detection and recognition.
Think of It Like This
Cones of personal space
Picture guests standing on a globe, each claiming a cone of personal space around their direction from the center. Softmax only asks that everyone stand somewhere inside their own cone. ArcFace widens every cone by a fixed angle, so guests must huddle near their cone's center line to fit. Tighter huddles with wider gaps between cones mean a newcomer is obviously inside one cone or none. The analogy stops at the enforcement: cones are re-carved every batch by gradient descent, not by a bouncer.
How It Actually Works
ArcFace L2-normalizes both the embedding and every class weight, so each logit is a cosine of the angle between the face and a class direction, scaled by s. Before the softmax it adds the margin m to the target class angle only, replacing cos(theta) with cos(theta + m). The network must therefore pull each face much closer to its class center to reach the same logit.
A worked margin
Take scale s = 64 and margin m = 0.5 radians (about 28.6 degrees). A face sits 20 degrees from its true class center: cos(20) is about 0.940, and with the margin the logit uses cos(48.6), about 0.661, giving 64 x 0.661 = 42.3. The nearest rival sits 60 degrees away with logit 64 x 0.5 = 32.0. The true class still wins by 10.3, but only because the face huddles close to its center; without the margin the same face would win by a lazy 28.2 and learn nothing about compactness.
Watch Out For
A margin so large that training collapses
Pushing m past about 0.5 radians on noisy labels makes the target logit nearly unreachable, and accuracy collapses instead of improving. The symptom is training loss that plateaus high while plain softmax trains fine. Fix it by starting near 0.35, cleaning label noise first, and only then raising the margin.
The Quick Version
- ArcFace normalizes embeddings and weights so logits are cosines on a hypersphere.
- Adding margin m to the target angle forces compact, well-separated identity clusters.
- It trains with ordinary class labels, avoiding the triplet mining that FaceNet needs.
- Scale s (often 64) controls how sharply angles convert into loss.
- Too large a margin on noisy data collapses training instead of helping.