DeepSORT Appearance Tracking
DeepSORT remembers what each person looks like, so tracks survive the short occlusions where pure motion matching hands out new identities.
Why Does This Exist?
SORT identifies targets by position alone, so anyone who steps behind a pillar comes back as a stranger. People, unlike blips, have persistent looks: jacket color, shape, gait texture. DeepSORT exists to give each track a visual memory, a person re-identification embedding, fused with motion in the matching cost. Occlusions that break SORT become routine re-findings. The paradigm is multi-object tracking.
Think of It Like This
A bouncer who remembers jackets
SORT is a bouncer who checks ticket stubs (positions) only. When the lights dip and stubs smudge, re-entry becomes guesswork. DeepSORT trains the bouncer to also remember jackets: a stub plus a red jacket re-admits the right person even after a blackout. The memory is a compact embedding, compared by cosine distance, and it ages as jackets change. The analogy stops at the gallery: the bouncer keeps one embedding per track, refreshed from recent sightings, not a photo album.
How It Actually Works
A small convolutional network pre-trained for person re-identification embeds every detection into a normalized feature vector. Association runs in cascaded stages: recently seen tracks match first (motion-gated plus appearance), then older coasting tracks get their chance, which prioritizes the living over the lost. Each track banks its recent embeddings and compares new detections by smallest cosine distance to the bank.
A worked re-finding
A track coasts behind a pillar with its Kalman prediction drifting, so the reappearing detection overlaps the prediction at IoU 0.2, far below any motion gate. But the detection embedding matches the track bank at cosine distance 0.09 (similarity 0.91), far inside the 0.2 appearance gate. Appearance overrules geometry, the track reclaims its identity, and the switch counter stays still where SORT would have minted a new ID.
Watch Out For
Embedding drift across camera changes
Re-identification embeddings trained on one camera network misjudge under new lighting and angles. The symptom is confident wrong matches after a camera swap, with distances that look healthy. Fix it by calibrating the appearance gate per deployment and fine-tuning the embedding network on local crops.
The Quick Version
- DeepSORT extends SORT with a re-identification embedding per detection.
- Cascaded matching serves young tracks before old coasting ones.
- Appearance rescues matches where motion prediction drifted during occlusion.
- Each track banks recent embeddings instead of trusting a single snapshot.
- Embedding gates must be recalibrated for new cameras and lighting.