ConvNeXt Modernized Convolutions
ConvNeXt rebuilds ResNet with transformer-era training and large depthwise kernels, proving tuned convolutions still match vision transformers on accuracy and speed.
Why Does This Exist?
Vision transformers overtook CNNs around 2021 and the field read that as convolutions being obsolete. But comparisons were unfair: transformers trained 300 epochs with AdamW, Mixup and heavy augmentation while ResNet baselines used 90-epoch SGD recipes from 2015.
ConvNeXt (Liu et al., 2022) modernizes the ResNet-50 backbone step by step toward Swin Transformer accuracy while staying fully convolutional. Patchify stem, depth ratios, large kernels, inverted bottlenecks and transformer training close the gap. This page covers those changes. Residual foundations live in residual networks and separable large kernels in MobileNet.
Think of It Like This
Restoring a classic car with modern parts
A 1990s sports car loses to modern rivals on old tires, old fuel and a worn suspension. Refit tires, fuel mapping, suspension and gearbox one part at a time and the classic chassis matches the newcomers, because the gap was maintenance rather than shape.
ResNet is the chassis. The 300-epoch recipe, AdamW, LayerNorm and 7x7 depthwise kernels are the refit. Where the analogy stops: some parts change the chassis itself, such as the patchify stem replacing the 7x7 stride-2 opener.
How It Actually Works
Changes apply in order. The stem becomes a 4x4 stride-4 patchify layer. Stage compute ratio moves from (3, 4, 6, 3) to (3, 3, 9, 3), matching Swin. Each block inverts: a 7x7 depthwise convolution first (only weights, so 4,704 at ), then LayerNorm, then 1x1 expand 4x with GELU, then 1x1 project back. Downsampling uses 2x2 stride-2 convolutions between stages.
Training carries half the gain: 300 epochs, AdamW, Mixup, CutMix, RandAugment and stochastic depth. ConvNeXt-Tiny reaches about 82.1% top-1 with 29 million parameters and 4.5 GFLOPs, against Swin-Tiny at 81.3% with 28 million and 4.5 GFLOPs. Same budget, convolution ahead.
Code
depthwise_7x7_c96 = 7 * 7 * 96print(depthwise_7x7_c96)# -> 4704
swin_t, convnext_t = 81.3, 82.1print(round(convnext_t - swin_t, 1))# -> 0.8Watch Out For
Using large kernels without the training recipe
A 7x7 depthwise kernel under a 90-epoch SGD schedule gains nothing and can lose accuracy. The symptom is a slower model scored as a failed idea. Adopt the 300-epoch AdamW recipe with Mixup and stochastic depth before judging the architecture.
Keeping BatchNorm inside the depthwise block
BatchNorm on small per-device batches destabilizes the wide depthwise maps. The symptom is noisy validation curves. Use LayerNorm after the depthwise step as the paper does, especially on multi-node runs.
The Quick Version
- ConvNeXt modernizes ResNet with a patchify stem, (3, 3, 9, 3) stages and 7x7 depthwise inverted blocks.
- A 7x7 depthwise layer at 96 channels costs only 4,704 weights, so large kernels stay cheap.
- Transformer-era training supplies half the gain, so 90-epoch schedules underrate the model.
- ConvNeXt-Tiny reaches about 82.1% top-1 against Swin-Tiny at 81.3% on equal budgets.
- Use LayerNorm in the blocks and the full 300-epoch recipe before comparing.