Skip to content
AI360Xpert
Beta

Two-Stream Networks for Video

Two-stream networks watch video with two eyes: one reads what frames look like, the other reads how pixels move, then they vote together.

Two-stream networks fuse looks with motion: spatial says swim 0.7, flow says run 0.6, half weight each lands swimming 0.50 to 0.40
Two-stream networks fuse looks with motion: spatial says swim 0.7, flow says run 0.6, half weight each lands swimming 0.50 to 0.40

Why Does This Exist?

Appearance lies: a pool suggests swimming even when nobody moves, and a track suggests running during a team photo. Single-frame models inherit the lie because they cannot see motion. Two-stream networks exist to give motion its own vote: a spatial CNN reads RGB frames while a temporal CNN reads stacks of optical flow fields, and late fusion decides. Either stream alone is biased; together they describe actions honestly. The broader context is video understanding.

Think of It Like This

A food critic with two reviewers

A restaurant critic sends a photographer (what does the plate look like) and a motion reviewer who watches the kitchen (how does the cooking move). The photographer loves garnish; the kitchen watcher loves technique. The final review averages both, so pretty-but-sloppy and ugly-but-skilled both land fairly. Two streams review video the same way. The analogy stops at the cost: the critic pays two salaries, and flow extraction plus a second CNN doubles the compute bill.

How It Actually Works

The spatial stream classifies single RGB frames with a standard image CNN. The temporal stream stacks horizontal and vertical flow fields over about ten frames and classifies the motion pattern with a twin CNN. Softmax scores fuse by averaging or a learned layer, and multi-clip sampling covers long videos. Training pre-trains both streams on image and flow data before joint fine-tuning, since flow fields look nothing like photos.

A worked fusion

A clip of a diver gets spatial scores of swimming 0.55, diving 0.30, other 0.15 (the pool dominates), while the motion stream scores diving 0.70, swimming 0.20, other 0.10 (the tuck-and-rotate pattern is unmistakable). Averaged fusion gives diving (0.30 + 0.70) / 2 = 0.50 versus swimming (0.55 + 0.20) / 2 = 0.375, so motion overrules the pool bias and the clip lands correctly.

Watch Out For

Paying for flow at serving time

Dense optical flow extraction can cost more than both CNNs combined, which kills real-time deployment. The symptom is a demo that runs at laboratory speed and a product that crawls. Fix it by distilling motion into fast approximations (motion vectors, learned flow) or by migrating to 3D CNNs that read RGB directly.

The Quick Version

  • Spatial stream reads appearance; temporal stream reads optical-flow motion.
  • Late fusion lets motion overrule scene-biased appearance and vice versa.
  • Flow stacks need their own pre-training since they look nothing like photos.
  • Flow extraction dominates serving cost and motivates faster successors.
  • The design pattern survives: separate what things look like from how they move.