Skip to content
AI360Xpert
Beta

MTCNN Cascaded Face Detection

MTCNN finds faces with three small networks in a row, where each stage throws away easy non-faces so the next stage can spend its effort on the hard cases.

MTCNN cascades three networks from thousands of proposals to one box plus five landmarks, and non-maximum suppression keeps box A at IoU 0.74 over B
MTCNN cascades three networks from thousands of proposals to one box plus five landmarks, and non-maximum suppression keeps box A at IoU 0.74 over B

Why Does This Exist?

Scanning every window of a photo with a big network is slow, and most windows contain no face at all. A single large detector wastes nearly all of its compute proving that walls, trees and skies are not faces. MTCNN (Multi-task Cascaded Convolutional Networks) exists to spend compute where it matters: a tiny network rejects the obvious background in milliseconds, and only the survivors reach the heavier stages. This page covers the three-stage cascade; the broader detect-then-recognize pipeline lives in face detection and recognition.

Think of It Like This

Three sieves with shrinking holes

Picture panning for gold with three sieves. The first sieve is coarse and fast: it dumps gravel by the bucketful and keeps anything that glints. The second sieve is finer and slower, working only on what the first one kept. The third is a jeweller's loupe that examines each remaining flake and notes its exact shape. MTCNN works the same way. P-Net is the coarse sieve, R-Net the fine one, and O-Net the loupe that also marks the eyes, nose and mouth corners. The analogy stops at the learning: sieves have fixed holes, while each MTCNN stage is a trained network.

How It Actually Works

MTCNN resizes the image into a pyramid so faces of any size pass through at a scale the networks expect, then runs three stages.

P-Net proposes, R-Net rejects, O-Net decides

P-Net is a fully convolutional network that slides over every scale and outputs face scores plus rough box offsets for thousands of windows. R-Net takes the surviving crops, rejects most false alarms, and tightens the boxes. O-Net examines the last survivors and outputs the final box plus five facial landmarks (both eye centers, nose tip, both mouth corners). Each stage trains on three tasks at once: face classification, box regression and landmark localization.

A worked suppression

Take two P-Net proposals on the same cheek. Box A is (100, 120, 60, 72) with score 0.92 and box B is (105, 125, 60, 72) with score 0.88, written as (x, y, width, height). The overlap in x runs from 105 to 160 (55 px) and in y from 125 to 192 (67 px), so the intersection is 3685 square px. Each box covers 4320 square px, giving a union of 4955 and an IoU of about 0.74. At a 0.7 threshold the lower-scoring box B is suppressed, and only A continues to R-Net.

Watch Out For

Tiny faces vanish without the pyramid

MTCNN only sees faces near the sizes its pyramid produces. If you skip scales to save time, faces under about 40 px across silently disappear while big faces still detect fine. The symptom is perfect results on portraits and zero detections on crowd photos. Fix it by building the pyramid with a small factor (around 0.709) so no size gap goes uncovered.

Reading landmarks off unaligned crops

The five O-Net landmarks sit in the coordinate frame of the refined box, not the original image. If you draw them on the raw photo without adding the box offset, every point lands in the wrong place and downstream alignment warps garbage. Always map landmarks back through the box position before using them.

The Quick Version

  • MTCNN is a three-stage cascade: P-Net proposes, R-Net rejects, O-Net finalizes boxes plus five landmarks.
  • An image pyramid lets fixed-size networks catch faces at any scale.
  • Each stage jointly learns classification, box regression and landmark positions.
  • Non-maximum suppression between stages removes duplicate proposals on the same face.
  • It is accurate on small faces but slower than single-shot detectors like RetinaFace.