Transposed Convolution
Transposed convolution learns how to upsample: each input pixel stamps a small kernel onto a larger canvas so decoders rebuild resolution step by step.
Why Does This Exist?
Encoders shrink images: 224 becomes 112, then 56, 28, 14. Segmentation and generation must climb back up to full resolution, one label per input pixel. Nearest-neighbour upsampling climbs back but learns nothing; every 2x2 block copies one value and boundaries stay blocky.
Transposed convolution makes the climb learnable. Each input value is multiplied by a small kernel and stamped onto the output canvas, with overlapping stamps summing where they land. Stride 2 with a 3x3 kernel doubles each side while the weights learn edge-aware interpolation. This page covers that upsampling pass. Downsampling lives in pooling and downsampling and size arithmetic in basic CNN operations.
Think of It Like This
Stamping tiles that overlap at the edges
Press an inked 3x3 sponge once per grid cell on a larger sheet, stepping two cells each time so stamps overlap by one row. Where two stamps overlap the ink adds up, and the learned sponge pattern decides whether edges come out sharp or soft.
Each input pixel is one press. The stride sets the step. The overlap is where learning happens, and also where artifacts start.
How It Actually Works
Output size for input , kernel , stride , padding and output padding is . A 14x14 map through , , gives , plus to reach exactly 28. That is the doubling path used in decoders.
Despite the old name, this is not a deconvolution that inverts anything. It is the gradient-shaped twin of a convolution: the forward pass broadcasts each input across several outputs instead of gathering many inputs into one. Uneven overlap is the known trap. With and some output cells receive two stamps and neighbours receive one, so a random kernel prints a checkerboard. The fixes are a kernel size divisible by the stride (4x4 at stride 2 overlaps evenly), initialization to bilinear weights, or replacing the layer with resize plus a regular convolution.
Code
def transposed_size(i, k=3, s=2, p=1, op=1): return (i - 1) * s - 2 * p + k + op
print(transposed_size(14))# -> 28
print(transposed_size(7))# -> 14Watch Out For
Checkerboard artifacts on smooth outputs
A 3x3 stride-2 layer with random weights prints alternating bright and dark cells on skies and walls. The symptom is a visible grid in generated images. Use a 4x4 kernel at stride 2, initialize from bilinear weights, or switch to nearest resize followed by a 3x3 convolution.
Off-by-one shapes that break skip connections
Transposed layers return 27 instead of 28 when output padding is forgotten, and concatenation with the encoder skip then crashes. Compute the size formula before training and assert the decoder map matches the skip map at every stage.
The Quick Version
- Transposed convolution upsamples by stamping a learned kernel per input pixel with overlapping sums.
- Size follows (I minus 1) times S minus 2P plus K plus OP, so 14 doubles to 28 with K3 S2 P1 OP1.
- It learns interpolation instead of copying, which sharpens segmentation and generation outputs.
- Kernel 3 at stride 2 overlaps unevenly and prints checkerboards; use kernel 4 or resize plus convolution.
- It is not an inverse of convolution, and skips demand exact shape matches.