Skip to content
AI360Xpert
Beta

R-CNN Object Detector

R-CNN runs a CNN on around two thousand proposed regions per image, which proved deep features beat hand-crafted ones but made detection far too slow for real time.

R-CNN extracts region proposals with selective search, then runs a CNN plus classifier on each of the two thousand regions independently.
R-CNN extracts region proposals with selective search, then runs a CNN plus classifier on each of the two thousand regions independently.

Why Does This Exist?

Before 2014, detectors ran hand-crafted HOG features through deformable part models. Girshick and colleagues asked what happens if each candidate region goes through a deep CNN instead. The answer, published in 2014, raised mean average precision on PASCAL VOC from around 33% to 58%: learned features crushed engineered ones. R-CNN matters today as the ancestor every later detector reacts to, and as the clearest illustration of the propose-then-classify idea.

Think of It Like This

Photocopying every suspect page

A clerk hunting a forged paragraph photocopies 2,000 candidate pages one by one, carries each to an expert, and waits for a verdict per page. Accurate, and absurd: the same document gets re-read 2,000 times.

R-CNN is that clerk. Selective search proposes the pages, the CNN expert reads each independently, and no computation is ever shared between overlapping regions.

How It Actually Works

The three-step pipeline

Propose. Selective search merges superpixels into roughly 2,000 category-independent region proposals per image, slow and CPU-bound, with no learning involved.

Describe. Each region is warped to 227 by 227 pixels and pushed through an AlexNet-style CNN independently. Overlapping regions recompute nearly identical convolutions thousands of times.

Decide. A set of linear SVMs scores each region's features per class, and a separate regressor tightens the boxes. Training is itself three disconnected stages: CNN fine-tuning, SVM fitting, regressor fitting.

The cost that killed it

Around 2,000 forward passes per image meant roughly 47 seconds per image on a 2014-era GPU. That single number motivated the entire successor line: Fast R-CNN shares one forward pass, Faster R-CNN learns the proposals, and single-shot detectors drop proposals entirely.

Code

# The R-CNN idea in miniature: one forward pass PER region (the bottleneck)import torch
cnn = torch.nn.Sequential(torch.nn.Conv2d(3, 64, 11, stride=4), torch.nn.ReLU(),                          torch.nn.AdaptiveAvgPool2d(1), torch.nn.Flatten())svm = torch.nn.Linear(64, 20)  # one-vs-rest scores per class
regions = [torch.randn(1, 3, 227, 227) for _ in range(2000)]scores = [svm(cnn(region)) for region in regions]  # 2000 CNN passes: the cost

Watch Out For

Copying the three-stage training today

R-CNN's split CNN, SVM and regressor training made sense before end-to-end detection losses existed. Rebuilding it now means three times the bugs for worse accuracy than any modern baseline. Study R-CNN for the idea; build on Faster R-CNN or a single-shot detector.

Warping regions destroys aspect clues

Forcing every region into a square warps tall pedestrians and wide cars identically, throwing away shape evidence. Later RoI methods pool from the feature map instead of warping pixels, which is one reason they score higher with less compute.

The Quick Version

  • R-CNN proved CNN features beat hand-crafted ones, jumping VOC mAP from 33% to 58%.
  • Pipeline: selective-search proposals, one CNN pass per region, SVM scoring plus regression.
  • Cost: around 2,000 forward passes and 47 seconds per image, which ruled out real time.
  • Its successors share computation: Fast R-CNN pools regions, Faster R-CNN learns proposals.