Siamese Networks Compared Pairs
A siamese network runs two inputs through twin networks with shared weights, learning similarity itself instead of memorizing categories.
Why Does This Exist?
Classifiers answer "which of my known classes is this," which is useless for signature checks, duplicate detection and one-shot learning where the classes at test time never appeared in training. The needed question is comparative: "are these two the same." Siamese networks exist to learn that comparison directly. Twin encoders with tied weights map both inputs into one space where distance decides, trained by contrastive loss or triplet loss. Trackers reuse the design in single-object tracking.
Think of It Like This
Twin judges with one rubric
Two judges score divers separately but trained from the same rubric book, so a 7.5 from either means the same thing. Their score gap, not the absolute numbers, decides whether two dives match. Siamese twins are the judges, shared weights are the rubric book, and the distance between embeddings is the score gap. The analogy stops at the symmetry requirement: real judges drift apart over time, while tied weights force the twins identical forever.
How It Actually Works
Both inputs pass through the same encoder, producing two embeddings. A distance layer (Euclidean or cosine) converts the pair into one similarity score, and the pair or triplet loss backpropagates through both branches, accumulating identical gradients into the shared weights. At test time the encoder runs once per input and distances against stored gallery embeddings decide.
A worked comparison
Twin encoders emit a = (0.6, 0.8) and b = (0.65, 0.76) for two signatures, both already unit length. Gaps are (0.05, -0.04), squared distance 0.0025 + 0.0016 = 0.0041, so Euclidean distance is about 0.064. Against a 0.5 operating threshold this pair is confidently same-writer. A forgery at c = (0.1, 0.99) gives gaps (0.5, 0.19), squared distance 0.25 + 0.0361 = 0.2861, distance about 0.535, which lands just over the line as different.
Watch Out For
Untied weights that silently fork
Copying one branch to start the other and then training both independently forks the twins: identical inputs start scoring differently. The symptom is similarity scores that drift with input order. Fix it by literally sharing one module instance between both paths, never two copies.
The Quick Version
- Siamese networks learn comparison, not classification, via twin shared-weight encoders.
- Distance between embeddings is the decision; thresholds come from validation pairs.
- Contrastive and triplet losses supply the training signal.
- One encoder serves both inputs at test time against a stored gallery.
- Weights must be shared by construction, never merely initialized equally.