CenterNet Detector
CenterNet detects each object as one center point on a heatmap, then reads off its size and position offset, skipping anchor boxes and their tuning.
Why Does This Exist?
Anchor pipelines spend their complexity budget before learning starts: template shapes, IoU thresholds, per-dataset clustering. CenterNet, from the 2019 Objects as Points paper, reframes detection as keypoint estimation, the same machinery that made pose estimation work. Each object becomes one point. The network predicts a heatmap of center likelihoods, reads peaks, and regresses width, height and a sub-pixel offset per peak. No anchors, no matching, minimal NMS pain.
Think of It Like This
Pinning butterflies by one pin each
A collector could frame every butterfly with a measured rectangle, or pin each with one pin through the thorax and note wingspan on the label. The pin locates, the label sizes.
CenterNet pins butterflies. The heatmap peak is the pin through the center, and the size regression is the wingspan on the label. One pin per insect, no frames to fit.
How It Actually Works
Heatmap, size and offset heads
Three heads share the backbone features. The heatmap head outputs one -channel map where ground-truth centers are splatted as Gaussians; training uses a penalty-reduced focal-style loss that forgives near-center misses. The size head regresses at each center, and the offset head corrects the rounding from feature-map stride, typically 4. Inference takes the top 100 heatmap peaks and decodes boxes as center plus size plus offset.
Why peaks replace NMS
Nearby points on the same object produce lower heatmap values than the true peak, so a 3 by 3 max-pooling peak extraction keeps only local maxima. Duplicate suppression falls out of the representation instead of a tuned IoU threshold, which is why CenterNet is simpler to deploy on new data.
Code
import torch
heatmap = torch.zeros(1, 3, 128, 128) # 3 classes, stride-4 mapheatmap[0, 1, 64, 60] = 0.95 # a peak: class 1 centered hereheatmap[0, 1, 64, 61] = 0.40 # its weaker neighbor
peaks = (heatmap == torch.nn.functional.max_pool2d(heatmap, 3, stride=1, padding=1))kept = (heatmap > 0.5) & peaks # only the 0.95 peak survivesWatch Out For
Center collisions in dense crowds
Two overlapping objects share nearly one center, and the heatmap holds a single merged peak. The symptom is undercounting in crowds where boxes would still separate. Dense scenes still favor box-based or query-based detectors over pure center points.
Downsampling stride eating small objects
At stride 4, a 12-pixel object is 3 heatmap cells wide and its Gaussian splat nearly vanishes. The symptom is missing tiny objects that anchor grids at fine levels catch. Raise input resolution or fuse finer features when small objects matter.
The Quick Version
- CenterNet represents each object as one center point on a class heatmap.
- Size and offset heads decode each peak into a full bounding box.
- Max-pooling peak extraction replaces most NMS tuning.
- Shared centers in crowds and stride effects on tiny objects are the limits.