Skip to content
AI360Xpert
Beta

FaceNet Embedding Recognition

FaceNet turns each face into a short list of numbers placed so that photos of the same person land close together and strangers land far apart.

A satisfied triplet teaches nothing: anchor-positive distance 0.0029 versus anchor-negative 0.53 clears margin 0.2, so loss equals 0
A satisfied triplet teaches nothing: anchor-positive distance 0.0029 versus anchor-negative 0.53 clears margin 0.2, so loss equals 0

Why Does This Exist?

A classifier with one output per person breaks the day a new person enrolls: adding stranger number 10,001 means retraining the whole network. Real systems enroll new faces constantly without retraining anything. FaceNet exists to make that possible by replacing classification with geometry. The network outputs an embedding, and identity becomes a nearest-neighbour lookup in that space. This page covers the embedding idea and its triplet training; the detect-align-match pipeline around it is in face detection and recognition.

Think of It Like This

A seating chart at a wedding

Imagine seating guests so that family members share tables and strangers sit far apart. When a new guest arrives, you do not redraw the whole chart: you just find the table where they fit. FaceNet lays out faces the same way. Training arranges the room, and enrollment seats one more guest without moving anyone else. The analogy stops at the dimensions: the room has hundreds of axes, not two, and closeness is squared Euclidean distance, not chairs.

How It Actually Works

An aligned face passes through a convolutional network (originally Inception-style, later ResNet-style) followed by L2 normalization, producing a unit-length embedding vector. Verification compares two embeddings with squared Euclidean distance against a threshold; identification searches the gallery for the nearest one.

Triplets carve the space

Training feeds triples: an anchor, a positive of the same person, and a negative of someone else. The triplet loss demands the anchor-positive distance stay smaller than the anchor-negative distance by at least a margin. Hard-example mining picks the triples that currently violate this, because easy triples teach nothing.

A worked triplet

Use 2D toy embeddings with squared distances and margin 0.2. Anchor a is (0.2, 0.3), positive p is (0.25, 0.32), negative n is (0.9, 0.1). The anchor-positive gaps are (0.05, 0.02), squared distance 0.0025 + 0.0004 = 0.0029. The anchor-negative gaps are (0.7, -0.2), squared distance 0.49 + 0.04 = 0.53. The loss is max(0, 0.0029 - 0.53 + 0.2), which is 0: this triplet is already satisfied and contributes no gradient.

Watch Out For

Training on easy triplets that teach nothing

Random triples are almost always easy, so the loss sits at zero while you pay for full training. The symptom is a model that converges fast and verifies poorly. The fix is online hard or semi-hard mining inside each batch: keep the negatives that currently sit closest to the anchor.

Comparing unnormalized embeddings

FaceNet thresholds assume unit-length vectors. If you skip L2 normalization at serving time, distances grow with vector magnitude instead of identity and your tuned threshold misfires. Always normalize before storing gallery embeddings and before comparing.

The Quick Version

  • FaceNet outputs an embedding, not a class label, so new people enroll without retraining.
  • Triplet loss pulls same-person pairs together and pushes strangers apart by a margin.
  • Hard-triplet mining inside each batch does the real teaching.
  • Identity is a distance threshold for verification or nearest neighbour for identification.
  • Embeddings must be L2-normalized before any comparison.