Skip to content
AI360Xpert
Beta

EfficientDet Detector

EfficientDet combines a weighted feature pyramid with compound scaling, so one knob trades accuracy against latency across phones to servers.

EfficientDet fuses pyramid features with learned weights and scales backbone, pyramid and heads together through one compound knob.
EfficientDet fuses pyramid features with learned weights and scales backbone, pyramid and heads together through one compound knob.

Why Does This Exist?

Detector tuning in 2019 meant hand surgery per deployment: swap the backbone, resize inputs, repeat pyramid layers until the latency budget fits. EfficientDet, published in 2020, ports EfficientNet's compound scaling to detection. One coefficient ϕ\phi scales backbone depth and width, pyramid depth, head depth and input resolution together, producing the D0 to D7 family from one design. D0 runs on mobile hardware; D7 chased the accuracy crown at a fraction of previous compute.

Think of It Like This

One thermostat for the whole house

Old houses have a heater dial, a window policy and a blanket strategy tuned separately per room. A modern thermostat takes one temperature setting and coordinates furnace, fans and vents together.

The ϕ\phi coefficient is that thermostat. Turn it from D0 to D4 and backbone, pyramid, heads and resolution all grow in fixed proportion, staying balanced instead of one part bottlenecking the rest.

How It Actually Works

BiFPN: weighted fusion

Standard pyramids add features equally, but not all inputs matter equally. BiFPN learns a weight per input and fuses with fast normalized weighting O=∑iwiIi/(∑jwj+ϵ)O = \sum_i w_i I_i / (\sum_j w_j + \epsilon), where each wiw_i passes through ReLU to stay non-negative. It also drops pyramid nodes with only one input edge, which contribute little, and repeats the bidirectional block as a layer. The same accuracy arrives with far fewer FLOPs than a naive repeated pyramid.

Compound scaling in practice

Scaling follows fixed ratios: resolution grows fastest since detection is resolution-hungry, while depth and width follow. The practical payoff is a smooth menu, D0 near 34ms latency for edge boxes, D4 for workstation GPUs, D7 for accuracy-first servers, with no redesign between rungs.

Code

import torch
def bifpn_fuse(features: list, weights: torch.Tensor, eps: float = 1e-4) -> torch.Tensor:    w = torch.relu(weights)    stacked = torch.stack(features, dim=0)    return (w.view(-1, 1, 1, 1) * stacked).sum(dim=0) / (w.sum() + eps)
fused = bifpn_fuse([torch.randn(1, 64, 32, 32) for _ in range(3)],                   torch.tensor([1.0, 0.5, 0.2]))

Watch Out For

Scaling resolution past your objects

Compound scaling assumes bigger inputs help, but upsampling tiny thumbnails invents no detail while burning latency quadratically. The symptom is slower inference with flat accuracy. Check that native object sizes justify the rung before climbing from D1 to D3.

BiFPN weights hiding a dead input

A learned weight can collapse to zero, silently amputating a pyramid level your small objects needed. The symptom is a scale-specific recall hole after training converges. Inspect the learned weights per level instead of treating fusion as a black box.

The Quick Version

  • EfficientDet scales backbone, pyramid, heads and resolution together with one coefficient.
  • BiFPN fuses pyramid levels with learned non-negative weights per input.
  • The D0 to D7 family spans mobile to server without redesign.
  • Resolution scaling only pays when native object sizes justify the pixels.