One-Stage Detectors
One-stage detectors predict every box and class in a single pass over a dense grid, skipping proposals entirely to reach video-rate speeds.
Why Does This Exist?
Two-stage detectors spend their budget proposing and refining a few hundred regions, which caps them near 5 to 10 frames per second. Self-driving cars and live surveillance need 30 to 60. One-stage detectors delete the proposal step: every grid cell of the feature map predicts boxes and classes directly in one forward pass. YOLO proved the concept in 2015, SSD made it multi-scale in 2016, and RetinaNet closed the accuracy gap in 2017 with focal loss.
Think of It Like This
One guard with a grid visor
A two-person security team posts a scout in the rafters plus an appraiser on the floor. A single guard wearing a grid visor scans the whole gallery in one sweeping glance, with every visor cell reporting what it sees simultaneously.
The sweep is faster and cheaper than coordinating two people, but crowded corners get less individual attention. That is the whole tradeoff: one glance for speed, two passes for thoroughness.
How It Actually Works
Dense predictions in one pass
The image goes through the backbone once. Each cell of each feature-map level predicts a fixed set of boxes: center offsets, sizes, an objectness or confidence score, and class probabilities. A 13 by 13 map with 3 anchors per cell yields 507 predictions; multi-scale heads across pyramid levels push totals past 10,000 boxes per image. Non-maximum suppression then keeps the best and discards the overlapping rest.
The imbalance problem and its fix
Fewer than 10 of those 10,000 boxes hold real objects, so plain cross-entropy drowns in easy background. Hard-negative mining caps the background ratio explicitly, while focal loss down-weights easy examples smoothly. Either fix is mandatory, not optional: without it the detector converges to predicting background everywhere.
Code
# Dense head sketch: every cell predicts boxes, objectness and classes at onceimport torch
features = torch.randn(1, 512, 13, 13) # one backbone passhead = torch.nn.Conv2d(512, 3 * (5 + 20), 1) # 3 anchors, 20 classes
out = head(features).permute(0, 2, 3, 1).reshape(1, -1, 25)boxes, objectness, classes = out[..., :4], out[..., 4:5], out[..., 5:]Watch Out For
Small objects falling between grid cells
Coarse grids assign tiny objects to cells dominated by background, so the box signal vanishes. The symptom is confident large-object detection beside invisible distant pedestrians. Use multi-scale heads or a feature pyramid so small objects meet fine-grained cells.
Skipping the imbalance fix
Training a dense head with plain cross-entropy looks fine for an epoch, then collapses to all-background as easy negatives dominate the gradient. The symptom is a loss that falls while recall sits at zero. Add focal loss or hard-negative mining from the first run.
The Quick Version
- One-stage detectors predict all boxes and classes in a single forward pass.
- Dense grids produce over 10,000 candidates per image, filtered by NMS.
- Foreground-background imbalance near 1000 to 1 must be countered with focal loss or mining.
- They reach 30 to 60 frames per second and trail two-stage models on small objects.