Skip to content
AI360Xpert
Beta

Top-K Accuracy

Top-k accuracy counts a prediction right when the true label lands anywhere in the model's k highest scores, forgiving confusion between near-duplicate classes.

Top-k accuracy checks whether the true label appears anywhere in the five highest scores, not only in first place.
Top-k accuracy checks whether the true label appears anywhere in the five highest scores, not only in first place.

Why Does This Exist?

ImageNet holds 1,000 classes including Siberian husky, Eskimo dog and malamute, three labels for what most people call the same dog. A model ranking the true label second is not wrong in any useful sense, yet top-1 accuracy marks it wrong. Top-k accuracy fixes the metric: the prediction counts as right when the true label appears anywhere in the top kk scores. AlexNet's 2012 breakthrough is quoted in top-5 error, 16% against the previous 26%, precisely because the benchmark designers knew exact-first-place was too strict.

Think of It Like This

A quiz that accepts the top five guesses

A quizmaster asking for a dog breed could demand one answer and fail everyone who says malamute instead of husky. Or they could accept any of your top five guesses, rewarding real knowledge while still failing a guess of cat.

Top-5 is that generous quizmaster. It forgives confusion inside a lookalike family and still punishes genuine ignorance.

How It Actually Works

Definition and a worked example

For each image, sort the KK class scores descending and check membership of the true label in the first kk. Top-1 is the special case k=1k = 1. Take scores dog 0.50, husky 0.30, cat 0.10, bird 0.06, car 0.04 with true label husky. Top-1 says wrong, dog won. Top-5 says right, husky sits second. Top-2 also says right.

When each k fits

Top-1 rules deployments with one action per prediction: a factory gate opens or it does not. Top-5 fits benchmarks and retrieval-style uses where a shortlist goes to a human or a second stage. Report both when classes overlap: the gap between them measures how much of your error is fine-grained confusion versus real failure. A 78% top-1 beside a 94% top-5 says the model knows the right neighborhood and needs part-level help, not new data.

Code

import torch
logits = torch.tensor([[2.0, 1.7, 0.1, -1.0, -2.0]])  # dog, husky, cat, bird, carlabels = torch.tensor([1])  # husky
top5 = logits.topk(5, dim=1).indicestop1_correct = (top5[:, :1] == labels.unsqueeze(1)).any(dim=1)top5_correct = (top5 == labels.unsqueeze(1)).any(dim=1)

Watch Out For

Top-5 hides a broken top-1 in production

A model with 95% top-5 and 70% top-1 looks great on a slide and fails at a gate that acts on the single winner. The symptom is benchmark pride beside deployment complaints. Always evaluate the kk your system actually consumes, usually 1.

Comparing top-k across different k or class counts

Top-5 over 1,000 classes is far stricter than top-5 over 10, where random guessing already scores 50%. The symptom is meaningless cross-dataset bragging. Quote kk with the class count, and compare top-1 to top-1.

The Quick Version

  • Top-k accuracy counts a prediction right when the true label ranks in the top kk scores.
  • Top-5 forgives near-duplicate-class confusion on benchmarks like ImageNet.
  • The top-1 to top-5 gap diagnoses fine-grained confusion versus genuine failure.
  • Deployments that act on one label must be judged by top-1, not top-5.