RetinaFace Joint Face Detector
RetinaFace finds faces in one forward pass, predicting the box, five landmarks and a 3D face shape together so each task sharpens the others.
Why Does This Exist?
Cascaded detectors like MTCNN are accurate but run three networks in sequence, which costs latency on every frame. Tiny faces in crowds were the other pain: a detector trained only on boxes kept missing them because a 16-pixel face carries almost no texture. RetinaFace exists to solve both at once with a single-shot design that learns extra face structure. By predicting landmarks and dense 3D vertices alongside each box, the network gets free supervision that anchors small faces. This page covers the joint single-shot design; the full pipeline around it is in face detection and recognition.
Think of It Like This
A portrait sketch artist
Imagine an artist who draws a stranger in one sitting without erasing. She blocks in the head oval first, then places the eyes, nose and mouth, then shades the cheekbones and jaw. Each layer constrains the next: once the eyes are placed, the head oval cannot drift. RetinaFace draws the same way. The box is the head oval, the five landmarks are the features, and the dense mesh is the shading. The analogy stops at the supervision: the artist uses taste, while RetinaFace learns every layer from labelled faces.
How It Actually Works
RetinaFace attaches several prediction heads to a feature pyramid backbone, so large faces are caught on coarse levels and tiny faces on fine ones.
One pass, four outputs
For every anchor on every pyramid level the network predicts a face score, box offsets, five landmark positions, and dense 3D vertex projections. All four losses train together, so landmark supervision pulls blurry small-face boxes into place even when the box signal alone is weak. Context modules widen the receptive field around each anchor, letting the network borrow surrounding hair, shoulders and background to confirm a doubtful face.
A worked landmark check
Suppose the ground-truth left eye sits at (206, 157) and the network predicts (205, 158). The error is the square root of (1 + 1), about 1.4 px. On a 24 px wide face that is under 6 percent of the face width, well inside the tolerance that keeps downstream alignment stable. A box-only detector with the same 1.4 px shift but no landmark loss would have nothing correcting a systematic drift across the whole crowd.
Watch Out For
Anchor scales that skip your face sizes
RetinaFace only detects faces that overlap its anchors. If your deployment sees 200 px portraits but the anchors top out near crowd sizes, big faces fragment into overlapping partial boxes. The symptom is double detections on close-ups. Fix it by matching anchor scales to the face sizes in your own images, not the defaults from the paper's crowd benchmark.
The Quick Version
- RetinaFace is a single-shot detector: one network predicts boxes, landmarks and dense mesh together.
- Landmark and mesh losses act as extra supervision that rescues tiny faces.
- A feature pyramid assigns each face size to the level with matching resolution.
- It runs faster than cascades because there are no sequential stages.
- Anchor scales must cover your real face sizes or detections fragment.