Skip to content
AI360Xpert
Beta

Multi-Label Image Classification

Multi-label image classification tags each photo with any number of labels at once, using one independent sigmoid per tag instead of a competing softmax.

A multi-label head scores every tag with its own sigmoid, so beach and sunset can both fire at once instead of competing.
A multi-label head scores every tag with its own sigmoid, so beach and sunset can both fire at once instead of competing.

Why Does This Exist?

Real photos carry several truths at once: a street scene is outdoor, daytime, crowded and urban. Multi-class classification forces one winner, so tagging systems built on softmax quietly drop every secondary tag. Multi-label classification is the separate architecture for that job: KK independent yes-or-no questions, one per tag, where circling one box takes nothing from the others.

Think of It Like This

A select-all-that-apply checklist

A single-choice ballot spoils if you mark twice. A property checklist asking about pool, parking and garden has no such rule: ticking pool takes nothing from parking, and a villa can hold all three.

Multi-label heads are the checklist. Each tag gets its own sigmoid judged on its own evidence, so a photo can be beach at 0.93 and sunset at 0.88 simultaneously, with neither score constraining the other.

How It Actually Works

Independent sigmoids and per-label loss

The head outputs one logit zkz_k per tag, each passed through its own sigmoid pk=1/(1+e−zk)p_k = 1 / (1 + e^{-z_k}). Nothing forces the KK probabilities to sum to 1. Training sums binary cross-entropy over every tag independently, so each output learns its own evidence threshold.

A worked example

Logits [2.0,1.0,0.1][2.0, 1.0, 0.1] for beach, sunset and crowd give sigmoids [0.881,0.731,0.525][0.881, 0.731, 0.525]. Compare with the softmax from the multi-class page, [0.659,0.242,0.099][0.659, 0.242, 0.099]: the same raw scores now let all three tags fire above a 0.5 threshold instead of crowning one winner. At inference each tag keeps its own threshold, tuned to its own cost of false alarms.

Code

import torch
head = torch.nn.Linear(2048, 3)  # beach, sunset, crowdcriterion = torch.nn.BCEWithLogitsLoss()
features = torch.randn(8, 2048)labels = torch.tensor([    [1., 1., 0.], [0., 1., 1.], [1., 0., 0.], [0., 0., 1.],    [1., 1., 1.], [0., 0., 0.], [1., 0., 1.], [0., 1., 0.],])
loss = criterion(head(features), labels)probs = torch.sigmoid(head(features))predicted = (probs > 0.5).long()

Watch Out For

Softmax habit on tagging data

Beginners reach for CrossEntropyLoss plus softmax on multi-tag data, and the network learns to pick the single strongest tag while starving the rest. The symptom is high precision on the top tag with collapsed recall everywhere else. Match the head to the task: sigmoid plus per-label binary cross-entropy.

One threshold for tags with different costs

Missing an explicit-content tag costs far more than missing a landscape tag, yet teams ship 0.5 everywhere. The symptom is either alert fatigue or missed violations. Tune each tag's threshold against its own precision-recall curve and its own real-world cost.

The Quick Version

  • Multi-label classification answers KK independent yes-or-no questions per photo.
  • Each tag uses its own sigmoid, so scores never compete for a fixed budget.
  • Training sums binary cross-entropy over all tags independently.
  • Every tag gets its own decision threshold tuned to its own costs.