Image Retrieval Visual Search
Image retrieval answers queries with pictures, ranking a gallery by embedding distance so the closest match to your photo or sketch floats to the top.
Why Does This Exist?
Shoppers photograph a chair and want that chair, tourists photograph a landmark and want its name, and moderators photograph a meme and want its origin. Keywords fail all three because the query is visual. Image retrieval exists to search by pixels: embed the query, rank the gallery by distance, return the neighbours. Metric learning trains the space; hard-negative mining keeps training honest.
Think of It Like This
A librarian who remembers every cover
A librarian who has memorized every book cover takes your torn dust jacket, walks the stacks, and returns with the closest matches armful-first. She compares wholes before details and knows which shelf sections to skip. Retrieval systems shelve gallery embeddings in index structures (inverted files, graph search) that skip the irrelevant millions the same way. The analogy stops at the memory: hers is visual and fuzzy, while the index is geometric and exact to the descriptor.
How It Actually Works
Backbones produce regional descriptors pooled into one compact global vector per image (GeM pooling, NetVLAD aggregation), trained with margin losses so same-object views cluster. At query time the vector searches an approximate nearest-neighbour index, then geometric verification re-ranks the shortlist by matching local features. Recall at K and mean average precision measure whether the true match lands in the returned handful.
A worked shortlist
A query chair searches 2 million products. The index returns 100 candidates in milliseconds; verification re-ranks them, placing the exact chair at rank 3 and two color variants at ranks 1 and 5. Recall at 10 scores 1.0 (truth in the handful), while average precision blends ranks 1, 3 and 5 into a single quality number that rewards putting the exact match first.
Watch Out For
Landmark models that memorize tourist angles
Training on landmark photos clusters the ten classic viewpoints and fails on night shots and close-ups. The symptom is perfect queries of postcard views with zero recall on user angles. Fix it by mining hard positives (same place, strange view) as aggressively as hard negatives.
The Quick Version
- Retrieval embeds the query and ranks the gallery by distance, no classifier involved.
- Global descriptors handle scale; local verification re-ranks the shortlist.
- Approximate indexes make million-image search interactive.
- Recall at K and mAP measure shortlist quality from different angles.
- Mine hard positives for viewpoint robustness, not just hard negatives.