Skip to content
AI360Xpert
Beta

RetinaNet Detector

RetinaNet pairs a feature pyramid with focal loss, which down-weights easy background boxes so the rare foreground objects steer training.

RetinaNet scores dense anchors across a feature pyramid while focal loss keeps easy background boxes from drowning the training signal.
RetinaNet scores dense anchors across a feature pyramid while focal loss keeps easy background boxes from drowning the training signal.

Why Does This Exist?

In 2017 one-stage detectors trailed two-stage accuracy by several points, and Lin and colleagues diagnosed why: dense grids produce around 100,000 candidates per image with single-digit foreground, so easy background boxes dominate the loss. RetinaNet keeps the one-stage speed and fixes the learning dynamics with focal loss, matching Faster R-CNN accuracy while staying single-shot. It proved the gap was a loss problem, not an architecture problem.

Think of It Like This

A teacher who stops grading easy quizzes

A teacher with 1,000 trivially correct quizzes and 10 tricky ones learns nothing by re-grading the easy pile. A smart teacher glances at the easy ones and spends grading time where students actually struggled.

Focal loss is that triage. Easy background boxes get near-zero weight automatically, so each gradient step is driven by the hard, ambiguous cases at the boundary.

How It Actually Works

Backbone plus two subnetworks

A ResNet with a feature pyramid supplies multi-scale maps P3 through P7. Two small fully convolutional subnetworks attach to every level: one predicts class scores for each anchor, the other predicts box offsets. Anchors cover 3 scales and 3 aspect ratios per location, and the two heads share weights across all pyramid levels.

Focal loss mechanics

Focal loss multiplies cross-entropy by (1−pt)γ(1 - p_t)^\gamma with γ=2\gamma = 2: FL(pt)=−(1−pt)γlog⁡(pt)FL(p_t) = -(1 - p_t)^\gamma \log(p_t), where ptp_t is the model's probability on the true class. A verified comparison shows the effect. At pt=0.9p_t = 0.9, cross-entropy is 0.1054 but focal loss is 0.0011, down roughly 100 times. At pt=0.3p_t = 0.3, cross-entropy is 1.2040 and focal loss 0.5900, down only about 2 times. Easy examples nearly vanish from the gradient while hard ones survive almost intact.

Code

import torch
def focal_loss(logits: torch.Tensor, targets: torch.Tensor, gamma: float = 2.0) -> torch.Tensor:    ce = torch.nn.functional.binary_cross_entropy_with_logits(logits, targets, reduction="none")    p_t = torch.exp(-ce)    return (((1 - p_t) ** gamma) * ce).mean()
logits = torch.tensor([2.2, -1.2])   # confident positive, confident negativetargets = torch.tensor([1., 0.])print(round(focal_loss(logits, targets).item(), 4))  # -> 0.0076

Watch Out For

Gamma left at 2 on every dataset

Gamma 2 fits COCO's imbalance; on milder ratios it over-suppresses and training stalls on merely medium examples. The symptom is a loss that flatlines early with under-confident scores. Treat gamma as a tuned knob, trying 0.5 to 2 against validation AP.

Initialization that drowns the first epochs

With 100,000-to-1 imbalance, random initial scores flood training with confident background mistakes before focal loss can help. RetinaNet initializes the final bias so each anchor starts near foreground probability 0.01. Skip that bias init and the first epochs diverge.

The Quick Version

  • RetinaNet combines a feature pyramid with focal loss in a one-stage design.
  • Focal loss scales cross-entropy by (1−pt)2(1 - p_t)^2, erasing easy background from the gradient.
  • Verified effect: 100x smaller loss on easy cases, only 2x smaller on hard ones.
  • It matched two-stage accuracy in 2017 while keeping single-shot speed.