SegNet Architecture
SegNet remembers where max-pooling picked each winner and reuses those pooling indices to upsample sharply without learning a decoder filter.
Why Does This Exist?
FCNs learn their upsampling with transposed convolutions, which costs parameters and can blur thin boundaries like poles and curbs. For road-scene parsing the boundary is the product: a two-pixel curb error moves the drivable edge.
SegNet (Badrinarayanan, Kendall, Cipolla, 2017) keeps the VGG-16 encoder but makes the decoder parameter-free at the upsampling step. Each max-pooling records which of the four pixels won; the decoder scatters features back to exactly those spots. This page covers index-guided upsampling. For learned feature copying instead, see U-Net.
Think of It Like This
Numbered coat-check tags
A cloakroom takes four coats per hook and hands back one tag per hook recording whose coat hangs in front. At pickup time the attendant does not guess: the tag says slot of hook . Pooling indices are those tags. The decoder hangs each feature back on its exact hook instead of spreading it across all four.
Where it stops: the tag remembers position but not appearance, so a trainable convolution still follows each unpooling to fill in texture.
How It Actually Works
The encoder mirrors VGG-16: thirteen convolutions in five blocks, each ending in max-pooling with stride . For every pooling window the argmax is stored.
Index-guided unpooling
The decoder mirrors the encoder. Its unpooling layer takes a map and scatters each value into a zero grid at the stored index, leaving the other three slots zero. A trainable convolution then densifies the sparse map. No transposed filter is learned for the spatial step, so the decoder holds far fewer parameters than FCN's.
Worked example
A encoder window holds with max at index (top-left). The stored index is . The decoder receives value for that cell and scatters it to in a block. The following convolution blends those spikes with neighbours, so a pole one pixel wide survives pooling instead of being averaged away.
Code
# Index unpooling for one 2x2 window.window = [3.0, 1.0, 0.5, 2.0]index = max(range(4), key=lambda i: window[i])out = [0.0, 0.0, 0.0, 0.0]out[index] = 4.5print((index, out))# -> (0, [4.5, 0.0, 0.0, 0.0])Watch Out For
Expecting indices to carry appearance
Indices carry position only. If you remove the convolutions after unpooling, masks turn into sparse spikes. Symptom: dotted predictions on textured regions. Fix: always keep the densifying block after every unpooling step.
Picking SegNet for scarce medical data
SegNet has no skip concatenation, so fine detail still attenuates through the bottleneck. On small medical sets U-Net with full feature skips beats it. Use SegNet for large road-scene sets where memory matters, not for thirty-scan clinical tasks.
The Quick Version
- SegNet stores the argmax index of every pooling window during encoding.
- The decoder scatters each feature back to its winning slot, then densifies with convolutions.
- Upsampling itself learns nothing, so the decoder stays small and edges stay crisp.
- It trails skip-concatenation designs on small medical datasets.
- Best fit: large road and indoor scene parsing where memory and sharp curbs matter.