SORT Tracking by Detection
SORT tracks many objects with a Kalman filter and optimal matching alone, proving how far plain motion geometry goes before appearance is needed.
Why Does This Exist?
Before SORT, multi-object trackers were slow batch optimizers that looked at whole videos offline. Live cameras cannot wait. SORT (Simple Online and Realtime Tracking) exists as the minimal online answer: trust a good detector, predict with constant-velocity physics, match optimally, and run at hundreds of frames per second. It sets the baseline every fancier tracker must beat. The paradigm around it is multi-object tracking; its appearance-based child is DeepSORT.
Think of It Like This
An air-traffic board with straight rulers
A controller plots each blip, lays a ruler along its last two positions, and slides it one step forward to guess the next blip. When guesses and blips pair up with least total sliding, assignments lock in. No photographs of planes, no airline logos: just positions and rulers. That is SORT. It works beautifully while flights hold course and fails when a plane circles behind a storm, because a ruler cannot identify a plane it cannot see.
How It Actually Works
Each track holds position, size, aspect and their velocities in a Kalman state. Every frame the filter predicts each track forward, IoU distances between predictions and new detections fill a cost matrix, and Hungarian matching pairs them. Matched tracks update with the detection; unmatched detections start trial tracks confirmed after consecutive hits; unmatched tracks coast briefly, then die.
A worked prediction and match
A track sits at x = 100 with velocity 10 px per frame, so the Kalman predict step places it at 110. The detector reports a box at 112, and the update settles the state near 111 depending on filter gain: prediction plus a fraction of the 2 px surprise. For assignment, costs of 0.3 and 0.2 on the diagonal versus 0.9 and 0.8 crossed give totals of 0.5 against 1.7, so identities stay uncrossed at minimum total cost.
Watch Out For
Long occlusions that motion cannot bridge
SORT has no memory of what targets look like, so a person hidden behind a bus for two seconds reappears as a stranger with a fresh ID. The symptom is sky-high ID-switch counts on crowded scenes despite good MOTA. Fix it by upgrading to appearance matching (DeepSORT) or by widening the coast window only when the camera is static.
The Quick Version
- SORT is the minimal online tracker: Kalman prediction plus Hungarian assignment on IoU cost.
- It assumes constant velocity and trusts the detector completely.
- Trial periods for births and grace windows for deaths tame flicker.
- Speed is the payoff: hundreds of frames per second on modest hardware.
- Occlusions and similar crossing targets need appearance features it does not have.