PSPNet Pyramid Parsing
PSPNet pools each location at four window sizes and fuses the results, so a pixel votes with knowledge of its room, its block, and the whole scene.
Why Does This Exist?
Local patches lie. A grey rectangle could be a road, a rooftop, or a desk depending on the scene around it, and FCN-style heads decide from a fixed window. PSPNet (Zhao et al., 2017) gives every pixel four levels of company: the whole scene, halves, thirds, and sixths. A pixel that looks like road but sits under office ceiling tiles gets outvoted by its global context.
This page covers the pyramid pooling module. For the dilated-filter alternative see DeepLab.
Think of It Like This
Four maps of the same neighbourhood
A lost courier carries four maps: the city overview, the district sheet, the block plan, and the street close-up. The close-up shows a grey square; the city map shows an office park, not a highway. The courier trusts the stack, not one sheet. Pyramid pooling staples those four maps to every pixel's local notes.
Where it stops: four fixed grids still quantize oddly shaped regions, so slivers smaller than a sixth of the map get few dedicated bins.
How It Actually Works
A ResNet backbone with dilated late stages outputs a map. The pyramid pooling module branches it four ways with adaptive average pooling to , , , and grids. Each branch is squeezed by a convolution to one quarter of the channels, bilinearly upsampled back to , and concatenated with the original map. A final convolution predicts classes. An auxiliary loss halfway up the backbone stabilizes training.
Worked example
Map size with channels. The branch pools each cell, squeezes to channels, and stretches back to . Concatenation yields channels per location. A grey pixel scoring road vs floor locally receives a global vote that scores indoor , flipping the fused decision to floor vs road .
Code
# Channel math of the pyramid pooling module.base, branches = 2048, 4squeezed = base // 4print(base + branches * squeezed)# -> 4096
# Cell width pooled by the 6x6 branch on a 64-wide map.print(round(64 / 6, 2))# -> 10.67Watch Out For
Upsampling the pyramid with mismatched alignment
Bilinear upsampling of the four branches must land exactly on the base map grid. Symptom: one-pixel ringing along straight walls after fusion. Fix: use aligned corners consistently and verify the upsampled branch shape equals the base map before concatenation.
Paying for four branches on tiny objects
The pyramid helps scene-scale ambiguity, not sub-pixel dots. On thin-defect data the extra channels add memory for little gain. Fix: profile with and without the module; for tiny-object sets prefer PointRend-style refinement instead.
The Quick Version
- PSPNet pools every location into , , , and context summaries.
- Each summary is squeezed, upsampled, and concatenated with the local map before prediction.
- Global context fixes road-versus-floor style confusions that fool local heads.
- An auxiliary mid-network loss keeps the deep ResNet training stable.
- Fixed grids still underserve very thin or oddly shaped regions.