Skip to content
AI360Xpert
Beta

Fast R-CNN Detector

Fast R-CNN runs the CNN once per image and pools each proposed region from the shared feature map, cutting detection time roughly tenfold over R-CNN.

Fast R-CNN computes one shared feature map per image, then pools every proposed region from it instead of running the CNN per region.
Fast R-CNN computes one shared feature map per image, then pools every proposed region from it instead of running the CNN per region.

Why Does This Exist?

R-CNN spent 47 seconds per image running 2,000 CNN passes over overlapping pixels. Girshick's 2015 fix keeps the external proposals but shares everything else: run the backbone once, project each proposal onto the resulting feature map, and pool a fixed-size crop per region. One forward pass replaces thousands, training collapses into a single multitask loss, and accuracy rises because the whole network learns together.

Think of It Like This

One survey photo, many magnifiers

R-CNN photographed each suspect house separately, flying the plane 2,000 times. Fast R-CNN flies once, prints one giant survey photo, and hands appraisers magnifiers to inspect any address on the shared print.

The survey photo is the single backbone pass. Each magnifier view is an RoI-pooled region. Same evidence, a tenth of the flying.

How It Actually Works

Shared features plus RoI pooling

The backbone converts the image into one feature map. Each selective-search proposal is projected onto that map and divided into a fixed grid, typically 7 by 7, with max-pooling per cell. Every region, whatever its pixel size, becomes the same fixed-length vector for the heads.

One multitask loss

Two sibling heads train jointly: a softmax classifier over C+1C + 1 classes and a per-class box regressor. The loss L=Lcls+λLregL = L_{cls} + \lambda L_{reg} replaces R-CNN's three disconnected training stages, and end-to-end learning lifts accuracy while cutting training time roughly threefold.

The remaining bottleneck is proposals themselves: selective search still costs about 2 seconds per image on CPU. That number is exactly what the region proposal network in Faster R-CNN was built to erase.

Code

# RoI pooling sketch: every region becomes a fixed 7x7 crop of shared featuresimport torch
feature_map = torch.randn(1, 512, 38, 50)   # one backbone passrois = torch.tensor([[0, 30., 40., 200., 220.],   # batch index + x1 y1 x2 y2                     [0, 120., 90., 300., 260.]])
pooled = torch.ops.torchvision.roi_pool(feature_map, rois, spatial_scale=1/16,                                        pooled_h=7, pooled_w=7)

Watch Out For

RoI pooling quantization misaligns small boxes

RoI pooling rounds floating-point region coordinates to feature-map integers twice, shifting tiny boxes by whole pixels. The symptom is a small-object accuracy gap that vanishes when you switch to RoIAlign's bilinear sampling. Use RoIAlign in any new build.

Proposals still dominate runtime

After the 10x speedup, selective search becomes the new 80% of runtime. Teams that profile only the GPU miss it because it runs on CPU. If proposals bottleneck you, that is the signal to move to Faster R-CNN's learned proposals.

The Quick Version

  • Fast R-CNN runs the backbone once and pools each region from the shared map.
  • RoI pooling converts any-size region into a fixed 7 by 7 feature crop.
  • One multitask loss trains classification and box regression end to end.
  • Speed rises roughly tenfold over R-CNN; selective-search proposals remain the bottleneck.