Binary Image Classification
Binary image classification answers one yes-or-no question per photo, such as defect or clean, with a single sigmoid output instead of a full multi-class head.
Why Does This Exist?
Most applied vision questions are yes-or-no: defect or clean, tumor or healthy, spam or not. Image classification covers the general machinery, but the two-class case deserves its own page because the head, the loss and the threshold decision all simplify, and the failure modes differ. There is no runner-up class to inspect, so calibration and the decision threshold carry the whole deployment.
Think of It Like This
A bouncer with one rule
A bouncer checking a single rule, over 21 or not, needs one glance at one ID field and one yes-or-no call. They do not rank the patron against every person in the city.
A binary head is that bouncer. One logit, one sigmoid, one threshold, typically 0.5. Everything else in the network just prepares the evidence for that single call.
How It Actually Works
Head, loss and threshold
The backbone produces a feature vector, and one linear unit outputs a logit . The predicted probability is , where is the sigmoid function mapping any real number into . Training minimizes binary cross-entropy , where is 1 for positive and 0 for negative.
At inference you compare against a threshold , usually 0.5. Raise to 0.8 and you trade recall for precision: fewer false alarms, more missed positives. That tradeoff is a product decision, not a model property, and it is the main reason binary heads stay deployed for years.
A worked example
A chest X-ray head outputs . Then . At the scan is flagged positive with 80% confidence. Move to 0.9 and the same scan flips to negative, which is correct behavior if false alarms page a radiologist at 3am.
Code
import torch
head = torch.nn.Linear(2048, 1)criterion = torch.nn.BCEWithLogitsLoss()
features = torch.randn(8, 2048) # backbone output for 8 imageslabels = torch.tensor([1., 0., 1., 1., 0., 0., 1., 0.]).unsqueeze(1)
logits = head(features)loss = criterion(logits, labels)probs = torch.sigmoid(logits)predicted = (probs > 0.5).long()Watch Out For
Accuracy lies under class imbalance
With 95% negatives, a model that always answers negative scores 95% accuracy while catching nothing. The symptom is a great accuracy number beside a recall near zero. Report precision, recall and the confusion matrix, and consider weighting the positive class in the loss.
Shipping the default 0.5 threshold
The 0.5 default assumes false positives and false negatives cost the same, which is rarely true in medicine or manufacturing. Tune on validation data against real costs before deployment, and re-tune it when the class ratio shifts.
The Quick Version
- Binary image classification answers one yes-or-no question per photo with a single sigmoid output.
- Training uses binary cross-entropy on the logit; inference thresholds the probability.
- The threshold trades precision against recall and is a product decision.
- Under imbalance, accuracy misleads, so report precision, recall and the confusion matrix.