Fully Convolutional Networks
Fully convolutional networks turned a classifier into a dense pixel labeller by swapping fully connected layers for 1x1 convolutions plus learned upsampling.
Why Does This Exist?
Image classifiers output one label per image, but semantic segmentation needs one label per pixel. The obvious fix of running a classifier on every sliding window repeats work thousands of times and still outputs blocky maps.
FCN (Long, Shelhamer, Darrell, 2015) showed the classifier already computes a spatial grid before its final layers. Convert those layers to convolutions, add learned upsampling, and the whole image is labelled in one forward pass. This page covers the conversion trick and the FCN-32s/16s/8s ladder. Neighbours cover the successors: SegNet for pooling indices and DeepLab for atrous context.
Think of It Like This
A rubber stamp grid instead of one stamp at a time
Imagine grading a wall of exam papers pinned in a grid. The old way picks up one paper, stamps it, puts it back, and repeats 10,000 times. FCN is a stamp as wide as the wall: one press grades every paper at once, then a second expanding press sharpens the blurry edges back to full size.
Where it stops: the giant stamp still sees coarsely, so fine edges need the skip fusions described below.
How It Actually Works
A VGG-style backbone downsamples a image by through five pooling stages, ending with a tensor. A fully connected layer reads that whole tensor at once, which is why it destroys spatial layout.
The convolutionalization trick
Replace each fully connected layer with a convolution with the same weights. A layer that mapped inputs to scores becomes filters of size . The output changes from a -vector to a score grid with classes. Nothing is relearned; the weights are reshaped, so ImageNet pretraining carries over.
Learned upsampling and the 32s/16s/8s ladder
A transposed convolution with stride stretches the coarse grid back to (FCN-32s), but edges are blobby. FCN-16s upsamples , adds the pool4 prediction map, then upsamples . FCN-8s repeats from pool3. Each skip re-injects a sharper map before the final stretch.
Worked example
Input . Backbone output is for PASCAL classes. A pixel at coarse position with winning class (person) and score vs runner-up covers a patch in the input, rows to , columns to . Bilinear initialization of the transposed filter spreads that smoothly, and the pool3 skip (at ) corrects the patch border by up to pixels where the person meets the background.
Code
# Coarse-grid geometry of FCN-32s on a 320x320 input.H, W, stride, classes = 320, 320, 32, 21gh, gw = H // stride, W // strideprint((gh, gw, classes))# -> (10, 10, 21)
# Input patch covered by coarse cell (4, 6).r, c = 4, 6print((r * stride, (r + 1) * stride, c * stride, (c + 1) * stride))# -> (128, 160, 192, 224)Watch Out For
Checkerboard artefacts from transposed convolutions
Stride- transposed convolutions with uneven kernel overlap paint a faint grid over the mask. Symptom: regular checkerboard flicker on flat regions like sky. Fix: initialize with bilinear weights, keep kernel size divisible by stride, or upsample with bilinear interpolation followed by a convolution.
Treating FCN-32s output as boundary truth
Coarse-only FCN-32s is off by many pixels at edges. Do not measure thin structures with it. Move to FCN-8s skips or a later encoder-decoder design before trusting boundaries.
The Quick Version
- FCN replaces fully connected layers with convolutions, keeping ImageNet weights and producing a spatial score grid.
- Transposed convolutions learn the upsampling from resolution back to full size.
- FCN-16s and FCN-8s fuse pool4 and pool3 maps to sharpen boundaries.
- One forward pass labels the whole image instead of thousands of sliding windows.
- Coarse-only output is blobby; skips or later decoders fix the edges.