1x1 Convolution
A 1x1 convolution mixes channels at each pixel without touching neighbours, so it squeezes wide layers down, expands them back up, and adds depth cheaply.
Why Does This Exist?
Wide layers are expensive. A 3x3 convolution from 256 channels to 256 channels holds weights, and stacking several of those burns compute fast. But much of that width is redundant: the layer mostly needs to recombine channel information, not re-read spatial patterns at full width.
The 1x1 convolution fixes this by separating channel mixing from spatial mixing. With kernel size 1 it reads no neighbours at all; at each pixel it applies one learned linear mix across channels, followed by a nonlinearity. Squeeze 256 channels to 64 with a 1x1, run the costly 3x3 at width 64, then expand back. This page covers that bottleneck pattern. Spatial scanning itself lives in convolutional layers and the arithmetic in basic CNN operations.
Think of It Like This
A mixing desk before the effects pedal
A band runs 256 microphone lines into a mixing desk that blends them down to 64 stems, runs those stems through one expensive effects pedal, then expands back to a full mix. The pedal hears every ingredient but processes far fewer lines.
The 1x1 layer is the desk. The 3x3 that follows is the pedal. Where the desk stops blending is where the analogy stops: the 1x1 adds its own nonlinearity, so it computes rather than just routing.
How It Actually Works
For input with channels, a 1x1 layer with filters holds weights of shape . At each position the output channel is . No neighbours enter. Spatial size is unchanged.
The bottleneck math is concrete. Direct 3x3 at full width costs weights. With a 64-channel bottleneck: the squeeze 1x1 costs , the 3x3 at width 64 costs , and the expand 1x1 costs . Total , about 8.5 times fewer, with two extra nonlinearities included.
Code
direct = 3 * 3 * 256 * 256print(direct)# -> 589824
squeeze = 256 * 64middle = 3 * 3 * 64 * 64expand = 64 * 256print(squeeze + middle + expand)# -> 69632
print(round(direct / (squeeze + middle + expand), 1))# -> 8.5Watch Out For
Expecting a 1x1 to add spatial context
A 1x1 kernel reads exactly one pixel, so stacking only 1x1 layers never grows the receptive field. The symptom is a network that classifies textures but cannot use shape or layout. Pair every bottleneck with a spatial kernel that actually looks at neighbours.
Squeezing channels until information chokes
Projecting 512 channels down to 8 destroys distinctions no later layer can recover. The symptom is training loss that stalls early while a wider bottleneck trains fine. Keep the squeeze ratio near 4 to 1 as in Inception blocks, and widen before assuming the idea failed.
The Quick Version
- A 1x1 convolution mixes channels at each pixel and reads no neighbours, leaving spatial size unchanged.
- Squeezing 256 channels to 64 before a 3x3 cuts that block from 589,824 weights to 69,632 with extra nonlinearities.
- Bottlenecks power Inception modules, ResNet blocks and pointwise steps of separable convolutions.
- A 1x1 never grows the receptive field, so it always pairs with a spatial kernel.
- Squeeze near 4 to 1; narrower projections choke information instead of compressing it.