Skip to content
AI360Xpert
Beta

Faster R-CNN Detector

Faster R-CNN replaces slow selective search with a Region Proposal Network that learns proposals on shared GPU features, unifying the whole detector into one network.

Faster R-CNN slides a Region Proposal Network over shared features to propose regions, then classifies each with a second-stage head.
Faster R-CNN slides a Region Proposal Network over shared features to propose regions, then classifies each with a second-stage head.

Why Does This Exist?

Fast R-CNN was fast except for selective search, which still burned 2 seconds per image on CPU. Ren and colleagues' 2015 insight: proposals can be learned by a small network sliding over the same GPU feature map the detector already computes. The Region Proposal Network costs almost nothing, runs in milliseconds, and the full system hits 5 frames per second while topping accuracy charts. It became the reference two-stage detector for a decade.

Think of It Like This

A scout who reads the same map

Fast R-CNN hired scouts who walked the terrain on foot while appraisers studied aerial photos. Faster R-CNN teaches the scouts to read the same aerial photo: they scan the shared map, circle promising blocks, and hand the marked map to the appraisers.

One map, two readers, no footwork. Proposals and classification finally share both the features and the GPU.

How It Actually Works

The RPN head

A 3 by 3 convolution slides over the backbone feature map. At each position, kk anchor boxes of different scales and aspect ratios, usually 9, each get an objectness score and four box offsets. Anchors overlapping a ground-truth box above IoU 0.7 train as positive, below 0.3 as background, and the rest are ignored. Top proposals pass through NMS down to about 300 survivors.

The second stage

Survivors are pooled with RoIAlign into fixed 7 by 7 crops, then split into a classifier over C+1C + 1 classes and a box regressor. The multitask loss adds the RPN's own classification and regression terms, so one backward pass trains proposals and decisions together. With a ResNet plus Feature Pyramid Network backbone, small-object recall jumps because each pyramid level proposes at its own scale.

Code

# Anchor decoding: offsets turn a static anchor into a proposalimport torch
anchors = torch.tensor([[150., 150., 50., 100.]])   # x_center, y_center, w, hoffsets = torch.tensor([[0.1, -0.2, 0.0, 0.1]])     # tx, ty, tw, th
xa, ya, wa, ha = anchors.unbind(1)tx, ty, tw, th = offsets.unbind(1)proposals = torch.stack([xa + tx * wa, ya + ty * ha,                         wa * torch.exp(tw), ha * torch.exp(th)], dim=1)

Watch Out For

Default anchors that ignore your object sizes

COCO-sized anchors miss tiny satellite ships and giant close-up defects alike. The symptom is confident detections on medium objects with vanished extremes. Cluster your own box sizes with k-means before training and set anchor scales to match.

Expecting video-rate speed

Five frames per second is research real-time, not deployment real-time. The symptom is a demo that stutters on live video. For 30-plus frames per second, switch families to a one-stage detector instead of pruning a two-stage one.

The Quick Version

  • Faster R-CNN learns proposals with an RPN on shared backbone features.
  • Anchors at each position get objectness scores plus box offsets, trimmed by NMS to 300.
  • RoIAlign plus joint multitask loss trains proposals and heads in one network.
  • It leads on accuracy, especially small objects with FPN, at about 5 frames per second.