Skip to content
AI360Xpert
Beta

Focal Loss for Segmentation

Focal loss multiplies cross-entropy by a factor that fades easy background pixels to near zero, leaving rare edge and lesion pixels in charge.

Focal loss turns down the volume on confident background pixels so a few uncertain edge pixels dominate the gradient.
Focal loss turns down the volume on confident background pixels so a few uncertain edge pixels dominate the gradient.

Why Does This Exist?

In dense masks 9090 percent or more of pixels are easy background the model masters in epoch one, yet plain cross-entropy keeps summing their small losses until they drown the few lesion and boundary pixels. Dice fixes this with overlap scoring; focal loss (Lin et al., 2017, adapted from detection to masks) fixes it by reweighting: multiply each pixel's cross-entropy by (1−pt)γ(1 - p_t)^\gamma with γ≈2\gamma \approx 2, where ptp_t is the predicted probability of the true class. Easy pixels fade; hard ones rule.

For the glossary entry see the focal loss definition. This page is the segmentation variant with per-pixel behaviour and numbers.

Think of It Like This

A coach who stops drilling mastered scales

A piano coach hears forty easy scales played right and two hard passages fumbled. Equal attention wastes the hour on scales. Focal loss hands the coach a volume knob wired to mastery: confident passages fade to a whisper, fumbled bars play loudly. Practice time flows to the failures.

Where it stops: the knob needs tuning. Too high a gamma mutes useful reinforcement and the easy passages drift back out of tune.

How It Actually Works

Per pixel with true-class probability ptp_t and balance factor αt\alpha_t:

Lfocal=−αt(1−pt)γlog⁡(pt)\mathcal{L}_{focal} = -\alpha_t (1 - p_t)^\gamma \log(p_t)

With γ=2\gamma = 2: a background pixel at pt=0.95p_t = 0.95 contributes (0.05)2=0.0025(0.05)^2 = 0.0025 of its cross-entropy, while an edge pixel at pt=0.3p_t = 0.3 keeps (0.7)2=0.49(0.7)^2 = 0.49. The easy pixel is downweighted roughly 200×200\times relative to the hard one before α\alpha balancing. α=0.25\alpha = 0.25 for foreground is the standard starting point from RetinaNet practice.

Worked example

Two pixels, γ=2\gamma = 2, no α\alpha. Easy: pt=0.9p_t = 0.9, cross-entropy −log⁡(0.9)≈0.1054-\log(0.9) \approx 0.1054, focal 0.01×0.1054≈0.00110.01 \times 0.1054 \approx 0.0011. Hard: pt=0.3p_t = 0.3, cross-entropy ≈1.2040\approx 1.2040, focal 0.49×1.2040≈0.59000.49 \times 1.2040 \approx 0.5900. Ratio hard-to-easy rises from about 11×11\times to about 536×536\times. That is why boundaries finally move.

Code

import math
def focal(pt: float, gamma: float = 2.0) -> float:    return -((1 - pt) ** gamma) * math.log(pt)
print(round(focal(0.9), 4), round(focal(0.3), 4))# -> (0.0011, 0.59)

Watch Out For

Gamma so high the model forgets easy regions

γ=5\gamma = 5 mutes background so hard that late-training drift reintroduces holes inside confident areas. Symptom: speckled interiors while edges look great. Fix: start at γ=2\gamma = 2, and combine with Dice if interiors destabilize.

Applying detection alpha blindly to masks

RetinaNet's α=0.25\alpha = 0.25 assumes box imbalance, not your organ ratio. Symptom: foreground recall stuck low despite focal loss. Fix: set α\alpha from your pixel frequencies and validate per-class recall, not just mean loss.

The Quick Version

  • Focal loss multiplies cross-entropy by (1−pt)γ(1 - p_t)^\gamma per pixel.
  • Confident background pixels fade by 100×100\times or more; hard pixels dominate.
  • Standard start: γ=2\gamma = 2, foreground α=0.25\alpha = 0.25, then tune to your frequencies.
  • Complements Dice: focal mines hard pixels, Dice guards global overlap.
  • Too much gamma destabilizes confident interiors.