Residual and Dense Networks
Instead of learning a full mapping from scratch, skip connections let layers learn only the residual difference, enabling networks hundreds of layers deep.
Why Does This Exist?
When deep neural networks were pushed past 20 layers in the mid-2010s, researchers observed a counterintuitive failure mode: the degradation problem. As network depth increased, training accuracy did not merely saturate—it dramatically worsened. A 56-layer plain convolutional network suffered higher training error and higher test error than a shallower 20-layer counterpart on CIFAR-10.
This failure was not caused by overfitting (since training error itself was higher) nor solely by vanishing gradients (which normalized initializations and Batch Normalization had largely stabilized). Instead, deep plain networks were mathematically incapable of learning simple identity mappings. If the optimal depth for a visual task is 20 layers, adding 36 additional layers should theoretically leave performance unchanged if those extra layers simply pass inputs through unmodified (). In practice, fitting an identity mapping using non-linear weight layers is extremely difficult for gradient descent to discover.
Residual Networks (ResNet, He et al., 2015) revolutionized deep learning by introducing identity shortcut connections. Instead of forcing layers to fit an underlying mapping , the network explicitly fits a residual function , outputting . This simple additive skip connection enabled networks to scale stably to 152 layers, and eventually over 1,000 layers. Subsequent architectures like ResNeXt (Xie et al., 2017) and DenseNet (Huang et al., 2017) extended this principle through multi-branch cardinality and dense feature-concatenation highways.
Think of It Like This
A manuscript editing pipeline with original carbon copies
Imagine writing a 150-page legal manuscript through a team of 150 sequential junior editors. In a plain network, each editor retypes the entire manuscript from memory based on the previous editor's draft. By editor 30, small omissions, phrasing shifts, and typos compound into garbled noise. If an editor has nothing useful to add, they struggle to produce an exact word-for-word copy of what they read.
A residual network clips the original draft to every editor's desk. Each editor is instructed only to mark red-pen corrections (the residual ) on an overlay sheet. If an editor has no improvements to make, they leave the overlay completely blank (), and the original document passes forward cleanly intact.
DenseNet takes this even further: instead of just receiving the draft from the immediately preceding editor, every single editor receives a complete folder containing the drafts written by every prior editor in the chain, allowing anyone to reference early raw sketches alongside recent revisions.
How It Actually Works
The ResNet Formulation and Gradient Highway
Formally, a residual building block computes:
where and are the input and output vectors, and is the residual mapping to be learned. When input and output channel dimensions differ (or when downsampling via stride), a linear projection aligns the shortcut:
For an arbitrary deep layer relative to an earlier layer , the forward state is:
During backpropagation, let denote the scalar loss function. Applying the multivariable chain rule yields:
The identity matrix creates an uninterrupted gradient highway. Even if the weight gradient term diminishes to zero, the gradient propagates directly back to layer without vanishing.
ResNeXt: Aggregated Residual Transformations and Cardinality
ResNeXt replaces the two-layer or three-layer bottleneck in ResNet with a set of parallel homogeneous transformations (termed cardinality):
Each path performs a low-dimensional bottleneck projection (). Rather than simply increasing network depth or channel width, increasing cardinality (typically ) increases representational capacity much more effectively without expanding FLOP count or parameter complexity. ResNeXt is mathematically equivalent to grouped convolutions with group count equal to .
DenseNet: Connecting Every Layer to Every Other Layer
DenseNet shifts from additive shortcuts () to channel concatenation (). For a block with layers, each layer receives the feature maps of all preceding layers as input:
If each function produces feature maps, the input to layer has channels, where is the network growth rate (typically small, e.g., ). Between dense blocks, transition layers apply a convolution and average pooling with a compression factor to control memory footprint.
Worked Example
Let us trace a ResNet bottleneck block taking an input tensor with stride :
- Input shape:
- Layer 1 (Reduce): Conv2d(256, 64, kernel=1, bias=False) BatchNorm2d ReLU
- Output shape:
- Parameters:
- Layer 2 (Spatial): Conv2d(64, 64, kernel=3, padding=1, bias=False) BatchNorm2d ReLU
- Output shape:
- Parameters:
- Layer 3 (Expand): Conv2d(64, 256, kernel=1, bias=False) BatchNorm2d
- Output shape:
- Parameters:
- Shortcut: Identity connection (since shapes match: , )
- Output:
- Final shape after ReLU:
Total parameters in this bottleneck: weights (excluding BatchNorm affine parameters). If we had used two plain convolutions with 256 channels, it would have cost parameters—over more parameters!
Code
import torchimport torch.nn as nn
class ResNetBottleneck(nn.Module): """Standard ResNet bottleneck block with additive identity shortcut.""" def __init__(self, in_channels: int, mid_channels: int, stride: int = 1) -> None: super().__init__() out_channels = mid_channels * 4 self.conv1 = nn.Conv2d(in_channels, mid_channels, kernel_size=1, bias=False) self.bn1 = nn.BatchNorm2d(mid_channels) self.conv2 = nn.Conv2d(mid_channels, mid_channels, kernel_size=3, stride=stride, padding=1, bias=False) self.bn2 = nn.BatchNorm2d(mid_channels) self.conv3 = nn.Conv2d(mid_channels, out_channels, kernel_size=1, bias=False) self.bn3 = nn.BatchNorm2d(out_channels) self.relu = nn.ReLU(inplace=True) # Shortcut projection if spatial dimension or channel count changes if stride != 1 or in_channels != out_channels: self.shortcut = nn.Sequential( nn.Conv2d(in_channels, out_channels, kernel_size=1, stride=stride, bias=False), nn.BatchNorm2d(out_channels) ) else: self.shortcut = nn.Identity()
def forward(self, x: torch.Tensor) -> torch.Tensor: identity = self.shortcut(x) out = self.relu(self.bn1(self.conv1(x))) out = self.relu(self.bn2(self.conv2(out))) out = self.bn3(self.conv3(out)) out += identity return self.relu(out)
# Verification with Batch=2, Channels=256, Height=56, Width=56x = torch.randn(2, 256, 56, 56)block = ResNetBottleneck(in_channels=256, mid_channels=64, stride=1)out = block(x)
print(out.shape)# -> torch.Size([2, 256, 56, 56])Watch Out For
In-place activation corruption on residual skip tensors
A notorious runtime error occurs when applying PyTorch's in-place nn.ReLU(inplace=True) directly before or during residual accumulation. If the tensor entering the block is modified in-place by an earlier operation, the identity shortcut will hold altered activation values by the time out += x executes.
Furthermore, during backpropagation, PyTorch requires the unmodified input to calculate gradients for the shortcut branch. If was modified in-place, the backward pass will crash with a RuntimeError: one of the variables needed for gradient computation has been modified by an inplace operation. Never perform in-place mutations on tensors that feed active residual branches.
The Quick Version
- ResNet resolves the network degradation problem by reformulating layer objectives to fit residual differences .
- The additive identity shortcut adds an identity matrix term to the backpropagation chain rule, preserving clean gradient flow across hundreds of layers.
- ResNeXt introduces cardinality (grouped parallel paths) as an architectural dimension that outperforms naive depth and width scaling.
- DenseNet replaces addition with channel concatenation, passing every layer's output to all subsequent layers for maximum feature reuse.
- Bottleneck designs () compress channels before expensive spatial convolutions, dramatically lowering computational cost.