EfficientDet Detector
EfficientDet combines a weighted feature pyramid with compound scaling, so one knob trades accuracy against latency across phones to servers.
Why Does This Exist?
Detector tuning in 2019 meant hand surgery per deployment: swap the backbone, resize inputs, repeat pyramid layers until the latency budget fits. EfficientDet, published in 2020, ports EfficientNet's compound scaling to detection. One coefficient scales backbone depth and width, pyramid depth, head depth and input resolution together, producing the D0 to D7 family from one design. D0 runs on mobile hardware; D7 chased the accuracy crown at a fraction of previous compute.
Think of It Like This
One thermostat for the whole house
Old houses have a heater dial, a window policy and a blanket strategy tuned separately per room. A modern thermostat takes one temperature setting and coordinates furnace, fans and vents together.
The coefficient is that thermostat. Turn it from D0 to D4 and backbone, pyramid, heads and resolution all grow in fixed proportion, staying balanced instead of one part bottlenecking the rest.
How It Actually Works
BiFPN: weighted fusion
Standard pyramids add features equally, but not all inputs matter equally. BiFPN learns a weight per input and fuses with fast normalized weighting , where each passes through ReLU to stay non-negative. It also drops pyramid nodes with only one input edge, which contribute little, and repeats the bidirectional block as a layer. The same accuracy arrives with far fewer FLOPs than a naive repeated pyramid.
Compound scaling in practice
Scaling follows fixed ratios: resolution grows fastest since detection is resolution-hungry, while depth and width follow. The practical payoff is a smooth menu, D0 near 34ms latency for edge boxes, D4 for workstation GPUs, D7 for accuracy-first servers, with no redesign between rungs.
Code
import torch
def bifpn_fuse(features: list, weights: torch.Tensor, eps: float = 1e-4) -> torch.Tensor: w = torch.relu(weights) stacked = torch.stack(features, dim=0) return (w.view(-1, 1, 1, 1) * stacked).sum(dim=0) / (w.sum() + eps)
fused = bifpn_fuse([torch.randn(1, 64, 32, 32) for _ in range(3)], torch.tensor([1.0, 0.5, 0.2]))Watch Out For
Scaling resolution past your objects
Compound scaling assumes bigger inputs help, but upsampling tiny thumbnails invents no detail while burning latency quadratically. The symptom is slower inference with flat accuracy. Check that native object sizes justify the rung before climbing from D1 to D3.
BiFPN weights hiding a dead input
A learned weight can collapse to zero, silently amputating a pyramid level your small objects needed. The symptom is a scale-specific recall hole after training converges. Inspect the learned weights per level instead of treating fusion as a black box.
The Quick Version
- EfficientDet scales backbone, pyramid, heads and resolution together with one coefficient.
- BiFPN fuses pyramid levels with learned non-negative weights per input.
- The D0 to D7 family spans mobile to server without redesign.
- Resolution scaling only pays when native object sizes justify the pixels.