Skip to content
AI360Xpert
Beta

Dice Loss for Segmentation

Dice loss scores the overlap between prediction and truth instead of counting pixels, so a tiny tumour still shouts loudly against a huge background.

Dice loss rewards the shared area between predicted and true masks relative to their combined size, not the count of background pixels.
Dice loss rewards the shared area between predicted and true masks relative to their combined size, not the count of background pixels.

Why Does This Exist?

Per-pixel cross-entropy counts every pixel equally, so a scan that is 9999 percent background lets the model score 9999 percent by predicting empty masks. Practitioners hit this in U-Net tumour work and in every small-lesion task: the loss looks great and the masks are blank. Dice loss (V-Net, Milletari et al., 2016) scores overlap 2∣X∩Y∣/(∣X∣+∣Y∣)2|X \cap Y| / (|X| + |Y|) and minimizes one minus that, so missing a 2020-pixel lesion hurts as much as missing a 20,00020{,}000-pixel organ.

This page is the overlap-loss base of the wave. Focal reweights hard pixels, Tversky splits false-positive and false-negative costs, and Lovasz optimizes IoU directly. General loss background lives in core ML and is out of scope here.

Think of It Like This

Grading tracings by overlap, not by paper

A teacher grading tracings of a coin on a poster board could count every matching square inch of board, which rewards leaving the sheet blank. Dice grading lays the tracing over the coin print and scores shared area against total inked area. A blank sheet scores zero no matter how big the board is.

Where it stops: overlap alone ignores calibration, so a model can be right-shaped but overconfident.

How It Actually Works

For predicted probabilities pip_i and binary truth gig_i over NN pixels, with smoothing ϵ\epsilon:

D=2∑ipigi+ϵ∑ipi2+∑igi2+ϵLDice=1−DD = \frac{2 \sum_i p_i g_i + \epsilon}{\sum_i p_i^2 + \sum_i g_i^2 + \epsilon} \qquad \mathcal{L}_{Dice} = 1 - D

Squares in the denominator (the V-Net form) downweight uncertain mid-range scores. Gradients flow through the intersection term, so shrinking a true region raises the loss even when the background is vast.

Worked example

Truth G=[1,1,0,0]G = [1, 1, 0, 0], prediction P=[0.8,0.7,0.1,0.2]P = [0.8, 0.7, 0.1, 0.2]. Numerator 2(0.8+0.7)=3.02(0.8 + 0.7) = 3.0. Denominator (0.64+0.49+0.01+0.04)+2.0=3.18(0.64 + 0.49 + 0.01 + 0.04) + 2.0 = 3.18. D=3.0/3.18≈0.9434D = 3.0 / 3.18 \approx 0.9434, loss ≈0.0566\approx 0.0566. A lazy all-zero prediction gives D=0D = 0, loss 1.01.0, while cross-entropy would call it 5050 percent accurate here and 9999 percent on real scans.

Code

# Soft Dice on the worked example.p = [0.8, 0.7, 0.1, 0.2]g = [1.0, 1.0, 0.0, 0.0]num = 2 * sum(pi * gi for pi, gi in zip(p, g))den = sum(pi * pi for pi in p) + sum(gi * gi for gi in g)print(round(1 - num / den, 4))# -> 0.0566

Watch Out For

Division wobble on empty slices

Slices with no foreground make the denominator tiny and gradients explode between batches. Symptom: loss spikes on background-only crops. Fix: keep ϵ≈10−5\epsilon \approx 10^{-5}, include foreground-bearing crops in every batch, or add a small cross-entropy term.

Using Dice alone for multi-class maps

Binary Dice per class ignores mutual exclusion. Symptom: overlapping class scores that sum past one. Fix: average per-class Dice with a softmax head, or pair Dice with cross-entropy.

The Quick Version

  • Dice loss is one minus the soft overlap 2∑pg/(∑p2+∑g2)2\sum pg / (\sum p^2 + \sum g^2).
  • Tiny foregrounds keep full gradient voice regardless of background size.
  • Blank predictions score the worst possible 1.01.0, not 9999 percent accuracy.
  • Needs smoothing and foreground-bearing batches for stability.
  • Pair with cross-entropy for multi-class or calibration-sensitive tasks.