Skip to content
AI360Xpert
Beta

Object Tracking Across Frames

Object tracking keeps one steady identity on each moving target across video frames, connecting lonely per-frame detections into life stories.

Tracking stitches per-frame detections into lasting identities, and overlap of IoU 0.747 clears the 0.5 success line
Tracking stitches per-frame detections into lasting identities, and overlap of IoU 0.747 clears the 0.5 success line

Why Does This Exist?

A detector sees each frame as a new world: the car in frame 1 and the car in frame 2 are unrelated boxes. Everything useful in video, counting cars through an intersection, following a player, flagging a left bag, needs the opposite: one identity carried through time. Object tracking exists to stitch detections into trajectories, predicting through gaps and resolving who is who when paths cross. Classical motion-first methods are covered in classical object tracking; this page maps the whole territory.

Think of It Like This

A teacher taking attendance on a field trip

A teacher counts heads at every stop (detection) but also knows each child by name across stops (tracking). When two children swap jackets behind the bus, headcounts stay right while names can flip. Single-object tracking is trailing one specific child by the hand. Multi-object tracking is keeping every name right through the jacket swaps, absences and newcomers. The analogy stops at the sensing: the teacher sees faces, while trackers match boxes, motion and appearance features.

How It Actually Works

Nearly all modern tracking is tracking-by-detection: a detector proposes boxes each frame, then an association stage links them to existing tracks. The field splits on how many targets are followed and what evidence links them.

The two branches

Single-object tracking starts from a box given in frame one and follows exactly that target, learning its appearance online. Multi-object tracking starts from nothing, detects everything, and solves assignment every frame: which detection continues which track, which starts a new one, which ends an old one. Motion models like the Kalman filter predict where each track should appear; appearance embeddings confirm it; the Hungarian algorithm settles competing claims optimally.

A worked overlap

A ground-truth box covers (50, 50, 100, 100) and the tracker predicts (60, 55, 100, 100) in (x, y, width, height). Overlap in x runs 60 to 150 (90 px) and in y 55 to 150 (95 px), giving 8550 square px of intersection. Each box covers 10000, so the union is 11450 and IoU is 0.747. Above the standard 0.5 success line, this frame counts as tracked, and the same IoU feeds both accuracy scores and association costs.

Watch Out For

Judging a tracker by its detector

Swapping in a stronger detector lifts tracking scores without changing any tracking logic, which fools teams into thinking association improved. The symptom is a leaderboard gain that vanishes the moment the detector is fixed across runs. Fix it by ablating with one frozen detector when comparing association methods.

The Quick Version

  • Tracking turns per-frame detections into persistent identities over time.
  • Single-object tracking follows one given target; multi-object tracking manages births, deaths and crossings.
  • Motion prediction, appearance matching and optimal assignment are the three ingredients.
  • IoU ties the pieces together as a success metric and a matching cost.
  • Occlusion, similar-looking targets and camera motion are the standard failure trio.