Two-Stage Detectors
Two-stage detectors first propose a few hundred regions that might hold objects, then classify and refine each one, trading speed for accuracy on small and crowded objects.
Why Does This Exist?
Sliding a classifier over 50,000 windows per image takes minutes per frame. Object detection needed a way to spend heavy computation only where objects might be. Two-stage detectors split the job: a cheap first stage throws away obvious background and proposes around 300 candidate regions, then an accurate second stage classifies each survivor and tightens its box. The lineage runs R-CNN to Fast R-CNN to Faster R-CNN, each generation moving more work onto shared features.
Think of It Like This
A scout team plus an appraisal team
An auction house clearing a mansion sends scouts first. They walk every room fast and tag items that might be valuable, ignoring empty walls and junk. Appraisers follow and spend real time only on tagged pieces, pricing each and recording details.
The proposal stage is the scouts: fast, high-recall, sloppy boxes. The classification stage is the appraisers: slow, precise, final. Two passes cost more than one glance, but nothing valuable gets skipped for speed.
How It Actually Works
Stage one: propose
A region proposal network slides over the backbone feature map, scores anchor boxes for objectness, and regresses rough boxes. Non-maximum suppression trims thousands of overlapping proposals to about 300 survivors per image.
Stage two: classify and refine
Each surviving region is cropped from the shared feature map into a fixed-size grid with RoIAlign, which samples with bilinear interpolation so box coordinates stay exact. Two heads finish the job: a classifier over classes including background, and a box regressor that nudges each box to its final coordinates. Training optimizes both heads jointly with a multitask loss.
The price is latency near 5 to 10 frames per second on older GPUs, against 60 plus for single-shot rivals. The reward is the best accuracy on small objects and dense crowds, which is why two-stage models still rule medical imaging and satellite analysis.
Code
# Stage 2 sketch: fixed-size region features into two headsimport torch
roi_features = torch.randn(300, 256, 7, 7) # 300 proposals, RoIAlign outputpooled = roi_features.flatten(start_dim=1) # -> (300, 12544)
classifier = torch.nn.Linear(12544, 81) # 80 classes + backgroundbox_head = torch.nn.Linear(12544, 80 * 4) # per-class box refinements
class_scores = classifier(pooled)box_deltas = box_head(pooled)Watch Out For
Paying two-stage latency for a one-stage problem
Highway vehicle counting at 60 frames per second gains nothing from 300 proposals and precise tiny-box recall. The symptom is a GPU-bound pipeline missing real-time deadlines. Match the family to the constraint: two-stage for accuracy-critical stills, one-stage for video-rate work.
Proposal recall caps everything downstream
The second stage only sees proposed regions, so a missed proposal is a missed object no classifier can recover. The symptom is vanished small objects with confident scores on everything found. Check proposal recall separately before tuning the classifier.
The Quick Version
- Two-stage detectors propose regions first, then classify and refine each survivor.
- Stage one favors recall with cheap objectness scoring; stage two favors precision.
- RoIAlign crops each region to fixed features without coordinate rounding errors.
- They lead on small-object accuracy and trail on speed, near 5 to 10 frames per second.