DeepLab v1 to v3 Plus
DeepLab keeps resolution high with atrous convolutions, sees many scales at once with ASPP, and sharpens edges with a small decoder in v3 Plus.
Why Does This Exist?
Downsampling by throws away the exact pixels semantic segmentation must label, and one fixed filter scale misses objects that range from a distant pedestrian to a near bus. The DeepLab line (Chen et al., 2015 to 2018) attacked both without exploding compute: hold the output stride at or with spaced-out filters, then probe every location at several fields of view at once.
This page traces v1 through v3 Plus as one arc. For the competing pyramid-pooling answer see PSPNet, and for object-context refinement see OCRNet.
Think of It Like This
A surveyor with a zoom rake
A surveyor must map pebbles and boulders in one pass. A normal rake touches adjacent soil only. An atrous rake skips teeth: same handle weight, wider reach. ASPP hands the surveyor four rakes at once with gaps of , and inches plus a whole-field glance, then merges the readings. v3 Plus adds a fine brush for the pebble outlines.
Where it stops: wider gaps sample sparsely, so very thin structures can fall between the teeth.
How It Actually Works
Atrous convolution
A filter with dilation rate inserts gaps between taps, reaching pixels across with weights. Rate covers ; rate covers . DeepLab replaces striding in the last blocks with dilation, so output stride stays (v1/v2) or (v3) instead of .
ASPP and the version ladder
v1 pairs atrous convolutions with a dense CRF that snaps labels to image edges. v2 introduces Atrous Spatial Pyramid Pooling: parallel branches at rates , , , plus a branch, concatenated and fused. v3 drops the CRF, adds image-level pooling, and batches normalization. v3 Plus adds a light decoder that upsamples the ASPP output and fuses it with low-level backbone features before the final stretch.
Worked example
Backbone at output stride on a image gives a map. An ASPP branch at rate has field on that map, which is input pixels across: enough to cover a bus. The rate- branch covers pixels for cars. The branch reads local texture. Concatenated, one location votes bus, car, and road texture together, and the v3 Plus decoder re-snaps the winning vote to the low-level edge map.
Code
# Atrous reach: a 3x3 kernel with rate r covers (2r+1) map pixels.def reach(rate: int) -> int: return 2 * rate + 1
rates = [1, 6, 12, 18]print([(r, reach(r), reach(r) * 16) for r in rates])# -> [(1, 3, 48), (6, 13, 208), (12, 25, 400), (18, 37, 592)]Watch Out For
Gridding artefacts at large rates
Rates above on a stride- map sample isolated pixels with dead gaps between taps. Symptom: striped misses on thin rails and wires. Fix: cap rates near the map size, add the image-pooling branch, or switch output stride to .
Shipping v1 CRF settings with v3
The dense CRF helped v1 but v3 Plus is accurate without it, and stale CRF sigmas blur the new sharp edges. Symptom: worse boundaries after adding post-processing. Fix: retune or drop the CRF when moving to v3 or v3 Plus.
The Quick Version
- Atrous convolution widens the receptive field with no extra weights and no extra downsampling.
- ASPP probes each location at rates , , plus image pooling to catch many object scales.
- v1 used a dense CRF; v2 added ASPP; v3 modernized training; v3 Plus added a light edge decoder.
- Output stride or keeps maps dense enough for boundaries.
- Over-large dilation rates cause gridding on thin structures.