Faster R-CNN Detector
Faster R-CNN replaces slow selective search with a Region Proposal Network that learns proposals on shared GPU features, unifying the whole detector into one network.
Why Does This Exist?
Fast R-CNN was fast except for selective search, which still burned 2 seconds per image on CPU. Ren and colleagues' 2015 insight: proposals can be learned by a small network sliding over the same GPU feature map the detector already computes. The Region Proposal Network costs almost nothing, runs in milliseconds, and the full system hits 5 frames per second while topping accuracy charts. It became the reference two-stage detector for a decade.
Think of It Like This
A scout who reads the same map
Fast R-CNN hired scouts who walked the terrain on foot while appraisers studied aerial photos. Faster R-CNN teaches the scouts to read the same aerial photo: they scan the shared map, circle promising blocks, and hand the marked map to the appraisers.
One map, two readers, no footwork. Proposals and classification finally share both the features and the GPU.
How It Actually Works
The RPN head
A 3 by 3 convolution slides over the backbone feature map. At each position, anchor boxes of different scales and aspect ratios, usually 9, each get an objectness score and four box offsets. Anchors overlapping a ground-truth box above IoU 0.7 train as positive, below 0.3 as background, and the rest are ignored. Top proposals pass through NMS down to about 300 survivors.
The second stage
Survivors are pooled with RoIAlign into fixed 7 by 7 crops, then split into a classifier over classes and a box regressor. The multitask loss adds the RPN's own classification and regression terms, so one backward pass trains proposals and decisions together. With a ResNet plus Feature Pyramid Network backbone, small-object recall jumps because each pyramid level proposes at its own scale.
Code
# Anchor decoding: offsets turn a static anchor into a proposalimport torch
anchors = torch.tensor([[150., 150., 50., 100.]]) # x_center, y_center, w, hoffsets = torch.tensor([[0.1, -0.2, 0.0, 0.1]]) # tx, ty, tw, th
xa, ya, wa, ha = anchors.unbind(1)tx, ty, tw, th = offsets.unbind(1)proposals = torch.stack([xa + tx * wa, ya + ty * ha, wa * torch.exp(tw), ha * torch.exp(th)], dim=1)Watch Out For
Default anchors that ignore your object sizes
COCO-sized anchors miss tiny satellite ships and giant close-up defects alike. The symptom is confident detections on medium objects with vanished extremes. Cluster your own box sizes with k-means before training and set anchor scales to match.
Expecting video-rate speed
Five frames per second is research real-time, not deployment real-time. The symptom is a demo that stutters on live video. For 30-plus frames per second, switch families to a one-stage detector instead of pruning a two-stage one.
The Quick Version
- Faster R-CNN learns proposals with an RPN on shared backbone features.
- Anchors at each position get objectness scores plus box offsets, trimmed by NMS to 300.
- RoIAlign plus joint multitask loss trains proposals and heads in one network.
- It leads on accuracy, especially small objects with FPN, at about 5 frames per second.