Skip to content
AI360Xpert
Beta

Cascade R-CNN Detector

Cascade R-CNN chains three detector stages with rising IoU thresholds, so each stage refines boxes to the quality the next stricter stage demands.

Cascade R-CNN passes boxes through three detector stages of rising IoU strictness, each refining the boxes for the next stage.
Cascade R-CNN passes boxes through three detector stages of rising IoU strictness, each refining the boxes for the next stage.

Why Does This Exist?

A detector trained at IoU threshold 0.5 accepts sloppy boxes, but raising the threshold to 0.7 starves training: too few proposals qualify as positive and the head overfits. This mismatch caps localization quality exactly where precision work needs it most. Cascade R-CNN, published in 2018, resolves the paradox with a sequence: stage one trains at 0.5 and outputs decent boxes, stage two trains at 0.6 on those improved boxes, stage three trains at 0.7 on further improved ones. Each stage sees enough positives at its own threshold because the previous stage upgraded the supply.

Think of It Like This

Three sanding passes, finer grit each time

A woodworker does not start with fine grit: coarse paper removes stock fast but scratches, medium smooths the scratches, fine polishes. Each pass prepares the surface the next grit needs.

Cascade stages are the grits. The 0.5 stage removes background coarsely, the 0.6 stage smooths the boxes, the 0.7 stage polishes positions. Skipping straight to fine grit on rough wood, like training one head at 0.7, just clogs the paper.

How It Actually Works

Resampling beats rethresholding

The key mechanism is resampling, not just three heads. Stage tt receives the regressed boxes of stage t−1t-1 as its input proposals. Because regression shifts boxes toward ground truth, the IoU distribution of the new proposals is higher, so threshold 0.6 or 0.7 still finds plenty of positives. Each head therefore trains on examples matched to its own strictness instead of drowning in negatives.

Inference and cost

At test time a proposal flows through all three stages, each refining the box before the final classifier scores it. High-IoU metrics respond most: COCO AP at strict thresholds rises 2 to 4 points over a single-head Faster R-CNN baseline. The cost is roughly three second-stage heads, so latency grows about 30 to 50 percent over Faster R-CNN.

Code

# Cascade inference: each stage refines the boxes the next stage scoresimport torch
boxes = torch.tensor([[100., 100., 200., 200.]])  # RPN proposalsfor head in (stage1_head, stage2_head, stage3_head):    deltas = head.regress(boxes)      # predict refinements at this strictness    boxes = head.apply_deltas(boxes, deltas)scores = stage3_head.classify(boxes)  # final verdict from the strictest head

Watch Out For

Judging it by AP at IoU 0.5 alone

The cascade's gains live at strict thresholds; AP50 barely moves while AP75 jumps. The symptom is a rewrite dismissed as pointless because the team read one column. Compare full COCO-style AP across thresholds before deciding.

Three heads tripling your overfitting

Each stage adds parameters hungry for positives, and small datasets cannot feed the strictest head. The symptom is stage-three scores oscillating while stage one stays stable. On small data, keep two stages or strengthen augmentation before going full cascade.

The Quick Version

  • Cascade R-CNN chains detector stages at IoU thresholds 0.5, 0.6 and 0.7.
  • Each stage's regression resamples better boxes, feeding positives to the next stricter head.
  • Gains concentrate at strict IoU metrics, worth 2 to 4 COCO AP points.
  • Cost is about 30 to 50 percent more latency than single-head Faster R-CNN.