Multi-Label Image Classification
Multi-label image classification tags each photo with any number of labels at once, using one independent sigmoid per tag instead of a competing softmax.
Why Does This Exist?
Real photos carry several truths at once: a street scene is outdoor, daytime, crowded and urban. Multi-class classification forces one winner, so tagging systems built on softmax quietly drop every secondary tag. Multi-label classification is the separate architecture for that job: independent yes-or-no questions, one per tag, where circling one box takes nothing from the others.
Think of It Like This
A select-all-that-apply checklist
A single-choice ballot spoils if you mark twice. A property checklist asking about pool, parking and garden has no such rule: ticking pool takes nothing from parking, and a villa can hold all three.
Multi-label heads are the checklist. Each tag gets its own sigmoid judged on its own evidence, so a photo can be beach at 0.93 and sunset at 0.88 simultaneously, with neither score constraining the other.
How It Actually Works
Independent sigmoids and per-label loss
The head outputs one logit per tag, each passed through its own sigmoid . Nothing forces the probabilities to sum to 1. Training sums binary cross-entropy over every tag independently, so each output learns its own evidence threshold.
A worked example
Logits for beach, sunset and crowd give sigmoids . Compare with the softmax from the multi-class page, : the same raw scores now let all three tags fire above a 0.5 threshold instead of crowning one winner. At inference each tag keeps its own threshold, tuned to its own cost of false alarms.
Code
import torch
head = torch.nn.Linear(2048, 3) # beach, sunset, crowdcriterion = torch.nn.BCEWithLogitsLoss()
features = torch.randn(8, 2048)labels = torch.tensor([ [1., 1., 0.], [0., 1., 1.], [1., 0., 0.], [0., 0., 1.], [1., 1., 1.], [0., 0., 0.], [1., 0., 1.], [0., 1., 0.],])
loss = criterion(head(features), labels)probs = torch.sigmoid(head(features))predicted = (probs > 0.5).long()Watch Out For
Softmax habit on tagging data
Beginners reach for CrossEntropyLoss plus softmax on multi-tag data, and the network learns to pick the single strongest tag while starving the rest. The symptom is high precision on the top tag with collapsed recall everywhere else. Match the head to the task: sigmoid plus per-label binary cross-entropy.
One threshold for tags with different costs
Missing an explicit-content tag costs far more than missing a landscape tag, yet teams ship 0.5 everywhere. The symptom is either alert fatigue or missed violations. Tune each tag's threshold against its own precision-recall curve and its own real-world cost.
The Quick Version
- Multi-label classification answers independent yes-or-no questions per photo.
- Each tag uses its own sigmoid, so scores never compete for a fixed budget.
- Training sums binary cross-entropy over all tags independently.
- Every tag gets its own decision threshold tuned to its own costs.