Skip to content
AI360Xpert
Beta

PSPNet Pyramid Parsing

PSPNet pools each location at four window sizes and fuses the results, so a pixel votes with knowledge of its room, its block, and the whole scene.

PSPNet summarizes the feature map in four pooled grids from whole-scene to fine blocks and fuses them before predicting.
PSPNet summarizes the feature map in four pooled grids from whole-scene to fine blocks and fuses them before predicting.

Why Does This Exist?

Local patches lie. A grey rectangle could be a road, a rooftop, or a desk depending on the scene around it, and FCN-style heads decide from a fixed window. PSPNet (Zhao et al., 2017) gives every pixel four levels of company: the whole scene, halves, thirds, and sixths. A pixel that looks like road but sits under office ceiling tiles gets outvoted by its global context.

This page covers the pyramid pooling module. For the dilated-filter alternative see DeepLab.

Think of It Like This

Four maps of the same neighbourhood

A lost courier carries four maps: the city overview, the district sheet, the block plan, and the street close-up. The close-up shows a grey square; the city map shows an office park, not a highway. The courier trusts the stack, not one sheet. Pyramid pooling staples those four maps to every pixel's local notes.

Where it stops: four fixed grids still quantize oddly shaped regions, so slivers smaller than a sixth of the map get few dedicated bins.

How It Actually Works

A ResNet backbone with dilated late stages outputs a H/8×W/8H/8 \times W/8 map. The pyramid pooling module branches it four ways with adaptive average pooling to 1×11 \times 1, 2×22 \times 2, 3×33 \times 3, and 6×66 \times 6 grids. Each branch is squeezed by a 1×11 \times 1 convolution to one quarter of the channels, bilinearly upsampled back to H/8×W/8H/8 \times W/8, and concatenated with the original map. A final convolution predicts KK classes. An auxiliary loss halfway up the backbone stabilizes training.

Worked example

Map size 64×6464 \times 64 with 20482048 channels. The 6×66 \times 6 branch pools each ≈10.7×10.7\approx 10.7 \times 10.7 cell, squeezes to 512512 channels, and stretches back to 64×6464 \times 64. Concatenation yields 2048+4×512=40962048 + 4 \times 512 = 4096 channels per location. A grey pixel scoring road 2.02.0 vs floor 1.91.9 locally receives a 1×11 \times 1 global vote that scores indoor +1.5+1.5, flipping the fused decision to floor 3.43.4 vs road 2.02.0.

Code

# Channel math of the pyramid pooling module.base, branches = 2048, 4squeezed = base // 4print(base + branches * squeezed)# -> 4096
# Cell width pooled by the 6x6 branch on a 64-wide map.print(round(64 / 6, 2))# -> 10.67

Watch Out For

Upsampling the pyramid with mismatched alignment

Bilinear upsampling of the four branches must land exactly on the base map grid. Symptom: one-pixel ringing along straight walls after fusion. Fix: use aligned corners consistently and verify the upsampled branch shape equals the base map before concatenation.

Paying for four branches on tiny objects

The pyramid helps scene-scale ambiguity, not sub-pixel dots. On thin-defect data the extra 20482048 channels add memory for little gain. Fix: profile with and without the module; for tiny-object sets prefer PointRend-style refinement instead.

The Quick Version

  • PSPNet pools every location into 1×11 \times 1, 2×22 \times 2, 3×33 \times 3, and 6×66 \times 6 context summaries.
  • Each summary is squeezed, upsampled, and concatenated with the local map before prediction.
  • Global context fixes road-versus-floor style confusions that fool local heads.
  • An auxiliary mid-network loss keeps the deep ResNet training stable.
  • Fixed grids still underserve very thin or oddly shaped regions.