Skip to content
AI360Xpert
Beta

SegNet Architecture

SegNet remembers where max-pooling picked each winner and reuses those pooling indices to upsample sharply without learning a decoder filter.

SegNet passes pooling indices from encoder to decoder so upsampling snaps each feature back to the pixel that won pooling.
SegNet passes pooling indices from encoder to decoder so upsampling snaps each feature back to the pixel that won pooling.

Why Does This Exist?

FCNs learn their upsampling with transposed convolutions, which costs parameters and can blur thin boundaries like poles and curbs. For road-scene parsing the boundary is the product: a two-pixel curb error moves the drivable edge.

SegNet (Badrinarayanan, Kendall, Cipolla, 2017) keeps the VGG-16 encoder but makes the decoder parameter-free at the upsampling step. Each 2×22 \times 2 max-pooling records which of the four pixels won; the decoder scatters features back to exactly those spots. This page covers index-guided upsampling. For learned feature copying instead, see U-Net.

Think of It Like This

Numbered coat-check tags

A cloakroom takes four coats per hook and hands back one tag per hook recording whose coat hangs in front. At pickup time the attendant does not guess: the tag says slot 33 of hook 1212. Pooling indices are those tags. The decoder hangs each feature back on its exact hook instead of spreading it across all four.

Where it stops: the tag remembers position but not appearance, so a trainable convolution still follows each unpooling to fill in texture.

How It Actually Works

The encoder mirrors VGG-16: thirteen 3×33 \times 3 convolutions in five blocks, each ending in 2×22 \times 2 max-pooling with stride 22. For every pooling window the argmax a∈{0,1,2,3}a \in \{0, 1, 2, 3\} is stored.

Index-guided unpooling

The decoder mirrors the encoder. Its unpooling layer takes a H/2×W/2H/2 \times W/2 map and scatters each value into a H×WH \times W zero grid at the stored index, leaving the other three slots zero. A trainable 3×33 \times 3 convolution then densifies the sparse map. No transposed filter is learned for the spatial step, so the decoder holds far fewer parameters than FCN's.

Worked example

A 2×22 \times 2 encoder window holds [3.0,1.0,0.5,2.0][3.0, 1.0, 0.5, 2.0] with max 3.03.0 at index 00 (top-left). The stored index is 00. The decoder receives value 4.54.5 for that cell and scatters it to [4.5,0,0,0][4.5, 0, 0, 0] in a 2×22 \times 2 block. The following 3×33 \times 3 convolution blends those spikes with neighbours, so a pole one pixel wide survives pooling instead of being averaged away.

Code

# Index unpooling for one 2x2 window.window = [3.0, 1.0, 0.5, 2.0]index = max(range(4), key=lambda i: window[i])out = [0.0, 0.0, 0.0, 0.0]out[index] = 4.5print((index, out))# -> (0, [4.5, 0.0, 0.0, 0.0])

Watch Out For

Expecting indices to carry appearance

Indices carry position only. If you remove the convolutions after unpooling, masks turn into sparse spikes. Symptom: dotted predictions on textured regions. Fix: always keep the densifying 3×33 \times 3 block after every unpooling step.

Picking SegNet for scarce medical data

SegNet has no skip concatenation, so fine detail still attenuates through the bottleneck. On small medical sets U-Net with full feature skips beats it. Use SegNet for large road-scene sets where memory matters, not for thirty-scan clinical tasks.

The Quick Version

  • SegNet stores the argmax index of every 2×22 \times 2 pooling window during encoding.
  • The decoder scatters each feature back to its winning slot, then densifies with convolutions.
  • Upsampling itself learns nothing, so the decoder stays small and edges stay crisp.
  • It trails skip-concatenation designs on small medical datasets.
  • Best fit: large road and indoor scene parsing where memory and sharp curbs matter.