Tracking Metrics MOTA IDF1 HOTA
Tracking metrics separate three questions: did you find everything, did you keep names right, and did you draw tight boxes.
Why Does This Exist?
A tracker can find every pedestrian yet scramble every name, or hold names perfectly while missing half the crowd. One score cannot praise and blame both halves. Tracking metrics exist as a trio that splits the credit: MOTA for detection hygiene, IDF1 for identity consistency, HOTA for their balance. Reading them together is the only honest way to compare SORT against DeepSORT.
Think of It Like This
Grading an exam with three sections
An exam has attendance (did you show up for every question), name tags (did each answer sit under the right heading) and neatness (were the boxes tight). MOTA grades attendance by counting blanks and scribbles. IDF1 grades the name tags globally. HOTA multiplies neatness by attendance per answer, then averages. A student can ace attendance while failing tags, exactly like a detector-heavy tracker with sparkling MOTA and scrambled identities.
How It Actually Works
MOTA equals 1 minus (misses + false positives + identity switches) over ground-truth boxes, and can go negative on terrible runs. IDF1 is the F1-score of identity matches: it maps each predicted identity to one true identity globally, then harmonizes identity precision and recall. HOTA averages, over matched pairs, the geometric mean of detection IoU and association accuracy, rewarding tight boxes that also keep their names.
A worked trio
Ten ground-truth boxes produce 1 miss, 1 false positive and 0 switches: MOTA = 1 - (1 + 1 + 0) / 10 = 0.80. The identities behind those boxes stay perfectly consistent, so IDF1 sits near 1.0 and exposes that association is fine while detection leaks. A rival run with 0 detection errors but 4 switches flips the story: MOTA 0.60 with IDF1 dragged down, correctly blaming association instead.
Watch Out For
Tuning to MOTA while identities burn
MOTA barely notices switches (one penalty each) beside misses and false alarms, so teams optimize detectors and declare tracking solved. The symptom is leaderboard MOTA with users complaining that counted people change numbers mid-store. Fix it by gating releases on IDF1 and switch counts, not MOTA alone.
The Quick Version
- MOTA = 1 - (misses + false alarms + switches) / ground truth; detection-heavy.
- IDF1 = identity F1 under global identity mapping; association-heavy.
- HOTA balances detection tightness with association accuracy per match.
- Report all three plus raw switch counts or single-metric tuning will mislead.
- Fix the detector to move MOTA; fix association to move IDF1.