Prototypical Networks Few-Shot
Prototypical networks classify with averages: each class becomes the mean of its few examples, and new inputs join the nearest average.
Why Does This Exist?
Fine-tuning a classifier on five images overfits instantly: hundreds of weights chase a handful of pixels. Yet humans learn a new animal from one picture. Prototypical networks exist as the simplest reliable recipe for this few-shot regime. No per-task fitting happens at all: embed the support examples, average each class, and classify queries by nearest mean. The whole method is one embedding function plus arithmetic.
Think of It Like This
Finding your team by its huddle
On a foggy field you cannot see faces, but each team huddles around its captain's average position. A lost player walks toward the nearest huddle center and is usually right. Prototypes are huddle centers: the mean of a class's few known players. The walk is Euclidean distance in embedding space. The analogy stops at the learning: players choose huddles, while the network learns an embedding where class members truly cluster around their mean.
How It Actually Works
Training mimics testing with episodes: sample a few classes, a few support images each, plus queries, then compute prototypes as support means and classify queries by softmax over negative squared distances. The episode loss backpropagates into the encoder, teaching an embedding where same-class items huddle. At deployment the identical routine runs once per new task with no gradient steps.
A worked 2-way 2-shot episode
Class A supports embed at (0.2, 0.3) and (0.4, 0.5), prototype (0.3, 0.4). Class B supports at (0.9, 0.1) and (1.1, 0.3), prototype (1.0, 0.2). A query at (0.35, 0.45) sits 0.0050 squared from A (gaps 0.05, 0.05) and 0.4850 from B (gaps -0.65, 0.25), so softmax over (-0.005, -0.485) assigns it to A with probability above 0.61. Two examples per class decided it.
Watch Out For
Outlier supports that drag the prototype
One mislabelled or blurry support example yanks the class mean toward the wrong region, and every query follows. The symptom is a whole class failing from a single bad shot. Fix it by screening supports with confidence scores or by using medoid-style robust averages when shots are noisy.
The Quick Version
- Each class prototype is the mean of its support embeddings.
- Queries join the nearest prototype by Euclidean distance.
- Episodic training matches the few-shot test setup exactly.
- No fine-tuning happens per task: the encoder plus averaging is the classifier.
- Clean supports matter enormously because means follow outliers.