Transformer-based Detectors
DETR replaces hand-crafted anchor boxes and NMS with a transformer encoder-decoder and bipartite Hungarian matching for true end-to-end detection.
Why Does This Exist?
Traditional object detectors such as Faster R-CNN and YOLO rely heavily on hand-crafted engineering heuristics:
- Anchor box design: Thousands of predefined bounding boxes with manually tuned spatial scales and aspect ratios that vary wildly between datasets.
- Rule-based label assignment: Complex IoU overlap thresholds (e.g., foreground, background) to match predictions to ground truth.
- Non-Maximum Suppression (NMS): Greedy post-processing steps that cannot be differentiated end-to-end and frequently fail in dense, crowded scenes.
These hand-tuned components made detection pipelines difficult to optimize and deploy cleanly across edge devices.
DETR (DEtection TRansformer, Carion et al., 2020) fundamentally re-imagined object detection as a direct set prediction problem. By pairing a standard CNN backbone with a transformer encoder-decoder and an end-to-end bipartite matching loss, DETR eliminated anchor boxes, rule-based assignment, and NMS entirely.
However, original DETR suffered from two major limitations: very slow training convergence (requiring 500 epochs to match Faster R-CNN) and weak performance on small objects due to quadratic computational complexity that prevented processing high-resolution feature maps. Deformable DETR (Zhu et al., 2020) resolved both bottlenecks using sparse, multi-scale deformable attention modules that sample only a few key reference points, cutting training time by (converging in 50 epochs).
Think of It Like This
A lottery ticket grid vs. an elite auction with numbered paddles
Imagine locating valuable antiques hidden throughout an estate sale.
Traditional anchor-based detectors (like Faster R-CNN or YOLO) carpet the entire house with 100,000 pre-printed lottery tickets: one for every square foot, every doorway, every tabletop. Many tickets will accidentally overlap the same antique chair. Afterwards, a cleaning crew has to rummage through the pile (NMS) to throw out redundant duplicates.
DETR operates like a private auction room with exactly numbered bidder paddles (the Object Queries). The transformer encoder reads the entire room inventory at once, taking in global context. Then, all 100 bidders raise their paddles simultaneously. Bidder 1 bids on the antique clock by the mantle; Bidder 2 bids on the oil painting above the sofa; Bidders 3 through 100 conclude there are no other antiques in the room and bid on "empty background" ().
An automated auction arbiter (the Hungarian matching algorithm) guarantees that each physical antique is awarded to at most one bidder. Because each antique has exactly one winner, there are no duplicates to delete and no NMS is needed.
How It Actually Works
The DETR Set Prediction Pipeline
- Backbone Feature Extraction: A CNN backbone (e.g., ResNet-50) processes an input image into a low-resolution feature map . A convolution reduces channels from to .
- Transformer Encoder: The spatial dimensions are flattened into a 1D sequence of length . Fixed 2D spatial sine-cosine positional encodings are added to each token. Standard multi-head self-attention models global pairwise spatial relationships across the entire image.
- Transformer Decoder & Object Queries: The decoder takes learnable positional embeddings called Object Queries (typically ). Through alternating self-attention and cross-attention over the encoder's memory, the queries iteratively interact and localize distinct physical objects.
- Prediction Feedforward Networks (FFNs): Each of the output embeddings passes through independent 3-layer MLPs:
- Linear classification head: outputs class probabilities over classes (including the null class ).
- Bounding box regression head: outputs normalized box coordinates .
Hungarian Bipartite Matching Loss
Let be the ground-truth set of objects, padded with null tokens up to size . We search for a permutation of elements with the lowest matching cost:
The pairwise matching cost between ground truth and prediction with predicted class probability and box is:
where the bounding box loss combines loss and Generalized IoU (GIoU) loss:
The optimal one-to-one assignment is solved efficiently in time using the Hungarian algorithm. Once matched, the network computes standard cross-entropy and box regression losses over the assigned pairs.
Deformable DETR: Multi-Scale Deformable Attention
Standard transformer self-attention over an image feature map calculates dot products for every single query point, resulting in prohibitive complexity.
Deformable DETR replaces full attention with multi-scale deformable attention. Given an input feature map and a 2D reference point , the deformable attention module queries only a small fixed set of sampling points (typically ) per attention head across feature levels:
where:
- indexes the attention head, indexes the feature level (FPN scale), and indexes the sampled key point.
- is a continuous 2D sampling offset predicted by a linear projection of the query embedding .
- is a normalized attention weight ().
- Bilinear interpolation evaluates at non-integer coordinates .
Because is small and fixed, complexity drops to , making multi-scale processing fast and practical.
Worked Example
Consider matching predictions against ground-truth objects (with 1 null token ).
Ground-truth objects:
- : Cat at
- : Dog at
- : Null token
Predictions:
- : Predicts Cat with at
- : Predicts Dog with at
- : Predicts Dog with at
Let simplified cost be: (for , 0 otherwise):
- For (Cat):
- With :
- With :
- With :
- For (Dog):
- With :
- With :
- With :
Cost matrix:
The Hungarian algorithm selects the optimal assignment: with total cost . Bipartite matching cleanly associates prediction 1 with the Cat, prediction 2 with the Dog, and leaves prediction 3 to be penalized as background!
Code
import torchfrom scipy.optimize import linear_sum_assignment
def hungarian_matching( pred_logits: torch.Tensor, pred_boxes: torch.Tensor, gt_labels: torch.Tensor, gt_boxes: torch.Tensor, cost_class: float = 1.0, cost_bbox: float = 5.0) -> list[tuple[int, int]]: """Solve optimal bipartite matching between DETR predictions and ground truth. pred_logits: shape (N, num_classes) pred_boxes: shape (N, 4) in (cx, cy, w, h) gt_labels: shape (M,) gt_boxes: shape (M, 4) """ n = pred_logits.shape[0] m = gt_labels.shape[0]
# Compute softmax class probabilities out_prob = pred_logits.softmax(-1) # Classification cost: -prob for ground truth class cost_class_matrix = -out_prob[:, gt_labels] # (N, M) # L1 Bounding box cost cost_bbox_matrix = torch.cdist(pred_boxes, gt_boxes, p=1) # (N, M) # Combined cost matrix (N, M) cost_matrix = cost_class * cost_class_matrix + cost_bbox * cost_bbox_matrix cost_matrix = cost_matrix.detach().cpu().numpy() # Solve Hungarian assignment pred_ind, gt_ind = linear_sum_assignment(cost_matrix) return list(zip(pred_ind, gt_ind))
# Verification with simulated inputspred_logits = torch.tensor([[5.0, -2.0], [-1.0, 4.0], [-3.0, -1.0]]) # N=3 queries, 2 classespred_boxes = torch.tensor([[0.2, 0.2, 0.4, 0.4], [0.7, 0.7, 0.3, 0.3], [0.1, 0.1, 0.2, 0.2]])gt_labels = torch.tensor([0, 1]) # M=2: object 0 (class 0), object 1 (class 1)gt_boxes = torch.tensor([[0.22, 0.19, 0.38, 0.41], [0.68, 0.71, 0.31, 0.29]])
matches = hungarian_matching(pred_logits, pred_boxes, gt_labels, gt_boxes)print(matches)# -> [(0, 0), (1, 1)]Watch Out For
Loss of spatial reference points and slow bipartite convergence
In original DETR, object queries learn static positional coordinates that struggle to locate small objects whose features change scale across images. This causes bipartite matching instability early in training, where different queries flip between the same objects across consecutive epochs, dragging training time out to 500 epochs.
Deformable DETR resolves this by using iterative bounding box refinement and explicit 2D reference points: each query explicitly predicts a 2D anchor center , and cross-attention only searches in the local neighborhood around that point. When implementing transformer detectors, always use explicit reference points or Two-Stage Deformable DETR proposals rather than unconstrained cross-attention queries.
The Quick Version
- DETR treats object detection as a direct set prediction task, removing hand-crafted anchor boxes and heuristic NMS post-processing.
- Transformer decoders query global image representations using learned Object Queries ( or ).
- Bipartite matching via the Hungarian algorithm finds an exact one-to-one assignment between predicted boxes and ground-truth objects.
- Deformable DETR replaces dense quadratic global attention with sparse multi-scale deformable attention sampling only points per head.
- Deformable DETR cuts training time from 500 epochs down to 50 epochs while dramatically improving small-object detection accuracy.