SSD Object Detector
SSD predicts boxes from several feature-map scales in one pass, so shallow layers catch small objects while deep layers catch large ones.
Why Does This Exist?
The original YOLO predicted from one coarse grid, so small objects blurred out. One-stage detectors needed multi-scale vision without the cost of an image pyramid. SSD, published in 2016, attaches predictors to six feature maps of shrinking resolution in a single pass: fine maps with small anchor boxes handle tiny objects, coarse maps with large anchors handle big ones. It ran 59 frames per second while matching Faster R-CNN accuracy on standard benchmarks.
Think of It Like This
Fishing with six net sizes at once
A fisher with one net catches one fish size. SSD casts six nets simultaneously: a fine mesh near the surface for minnows, progressively wider meshes deeper for bigger fish.
Each feature-map scale is one net, and each anchor size is its mesh. One cast, the whole size range covered.
How It Actually Works
Multi-scale heads and matching
Extra convolutional layers extend the backbone into a descending staircase, for example 38 by 38 down to 1 by 1. Each level tiles anchors sized to its resolution and predicts class scores plus box offsets directly. During training, each ground-truth box matches the anchor with the best IoU above 0.5, plus the single best anchor regardless of threshold, guaranteeing every object supervises at least one predictor.
Hard-negative mining
With tens of thousands of anchors per image, negatives outnumber positives roughly 100 to 1. SSD sorts negative boxes by confidence loss and keeps only the top three per positive, a 3 to 1 cap that keeps the gradient honest. This explicit mining predates focal loss and solves the same drowning problem.
Code
# Multi-scale SSD sketch: one head per feature level, anchors sized per levelimport torch
levels = [torch.randn(1, 512, s, s) for s in (38, 19, 10, 5, 3, 1)]anchor_counts = [4, 6, 6, 6, 4, 4]
for features, n_anchors in zip(levels, anchor_counts): head = torch.nn.Conv2d(512, n_anchors * (4 + 21), 3, padding=1) out = head(features) # class scores + box offsets at this scaleWatch Out For
Anchor scales copied blindly to a new dataset
Default scales assume COCO-sized objects. On aerial imagery where everything is tiny, most ground-truth boxes match nothing and recall collapses at the matching step, before any learning happens. Verify match coverage per scale on your own boxes first.
Expecting RetinaNet-level dense scenes
SSD predates focal loss and feature pyramids, so crowded small-object scenes still favor its successors. The symptom is merged or dropped boxes in crowds that RetinaNet separates. Pick SSD for speed on moderate scenes, not for dense crowds.
The Quick Version
- SSD predicts from six feature-map scales in one forward pass at 59 frames per second.
- Fine maps with small anchors catch small objects; coarse maps catch large ones.
- Ground-truth boxes match anchors above IoU 0.5, with hard-negative mining capped at 3 to 1.
- Anchor scales must be re-tuned per dataset or small objects match nothing.