ShuffleNet Channel Shuffle
ShuffleNet runs convolution in isolated channel groups and then shuffles channels between groups, cutting cost while keeping information flowing across the whole layer.
Why Does This Exist?
Pointwise 1x1 layers dominate mobile cost: in separable stacks they hold most of the weights. Grouped convolution divides that cost by the group count , but isolated groups never exchange information, so stacking grouped layers builds independent sub-networks that share nothing.
ShuffleNet (Zhang et al., 2017) fixes the isolation with a channel shuffle: reshape, transpose and flatten so each next group receives channels from every previous group. Cost stays divided by while information mixes fully. This page covers grouped pointwise layers plus shuffle. Separable spatial steps live in MobileNet and the 1x1 mechanism in 1x1 convolution.
Think of It Like This
Study groups that remix every round
Students work in fixed groups of four, then everyone is dealt into new groups mixing one member from each old group. Each round stays small and cheap, but after two rounds every student has worked with students from every original group.
Grouped convolution is the fixed groups. The shuffle is the redeal. Where the analogy stops: the redeal pattern is a fixed transpose, not a learned routing.
How It Actually Works
A 1x1 from 240 to 240 channels costs weights. With groups each group maps 30 to 30 channels, costing , an 8-fold cut. The shuffle then transposes the layout so output group holds input channel from every input group.
ShuffleNetV2 (2018) adds the deployment lesson: FLOPs mislead because memory access dominates. It splits channels instead of grouping at the unit input, keeps one branch untouched, processes the other with depthwise plus pointwise steps, concatenates and shuffles. Equal channel widths minimize access cost, which is why V2 beats V1 at identical FLOPs.
Code
full = 240 * 240grouped = 8 * 30 * 30print(full, grouped)# -> (57600, 7200)
print(full // grouped)# -> 8Watch Out For
Stacking grouped layers without a shuffle between them
Two grouped 1x1 layers with the same grouping never mix information across groups. The symptom is accuracy stuck near a narrower network despite the parameter count. Place a shuffle, or change group counts, between every pair of grouped layers.
Optimizing FLOPs while ignoring memory access
Eight groups minimize FLOPs but fragment memory reads. The symptom is a low-FLOP model that runs slower than a wider one. Follow the V2 rules: equal widths, fewer groups at large maps, and measure latency on the target chip.
The Quick Version
- Grouped 1x1 layers divide pointwise cost by the group count, so 240 channels at 8 groups cost 7,200 weights not 57,600.
- Channel shuffle remixes groups so information flows across the whole layer despite the split.
- ShuffleNetV2 adds split-concat units and equal channel widths to cut memory access, not just FLOPs.
- Stacking grouped layers without shuffling builds isolated sub-networks that share nothing.
- Measure on-device latency because fragmentation can erase FLOP savings.