Skip to content
AI360Xpert
Beta

SlowFast Networks for Video

SlowFast watches video at two speeds: a slow path studies sharp appearances while a fast path catches quick motion, then they compare notes.

SlowFast pairs a slow full detail path with a fast low channel path: 4 big frames plus 32 thin frames, alpha 8 beta 1/8
SlowFast pairs a slow full detail path with a fast low channel path: 4 big frames plus 32 thin frames, alpha 8 beta 1/8

Why Does This Exist?

Uniform frame sampling wastes resolution on still backgrounds and starves fast motion of temporal detail: one rate cannot serve both. Two-stream networks split appearance from motion but pay for optical flow. SlowFast exists as the RGB-only answer with two frame rates inside one network. A slow path samples sparsely at full detail for semantics; a fast path samples densely at reduced channels for tempo. Lateral connections fuse them at several depths. The task context is action recognition.

Think of It Like This

A painter and a drummer watching a dance

A painter sketches one pose per bar in full detail while a drummer taps every beat with eyes half closed. The painter knows who dances; the drummer knows the rhythm. Between bars the drummer's taps annotate the painter's canvas with tempo marks. SlowFast pairs the same two observers: the slow path paints semantics, the fast path drums motion, and lateral links annotate across. The analogy stops at the channels: the drummer uses fewer neurons per tap (about one-eighth the channels) because rhythm needs less detail than identity.

How It Actually Works

The slow path takes roughly 8 frames across the clip through a full 3D ResNet; the fast path takes roughly 32 frames (4x the rate) through a thin network with about one-eighth the channels. Time-strided convolutions in the fast path preserve tempo while lateral concatenations inject motion features into the slow path at each stage. One classifier head reads the fused representation.

A worked budget

A 64-frame window feeds 8 frames to the slow path and 32 to the fast path. If a slow frame costs 1.0 unit of compute, a fast frame at one-eighth channels costs about 0.125, so the fast path adds 32 x 0.125 = 4.0 units against the slow path's 8.0: tempo costs half as much as semantics while quadrupling the frame rate. That ratio is the design's whole economy.

Watch Out For

Slow-only ablations that hide tempo failures

The slow path alone scores well on scene-biased benchmarks, tempting teams to delete the fast path for speed. The symptom is a lean model that cannot tell fast claps from slow waves. Fix it by benchmarking on motion-sensitive splits (something-something style) where tempo decides, before amputating.

The Quick Version

  • Slow path: sparse frames, full channels, semantics. Fast path: dense frames, thin channels, tempo.
  • Lateral connections fuse motion into semantics at multiple depths.
  • No optical flow needed: tempo comes from raw RGB at high frame rate.
  • Fast path costs a fraction per frame by design, keeping the budget sane.
  • Verify on motion-sensitive data or scene bias will hide tempo blindness.