Pretrained Vision Models
A pretrained vision model ships with backbone weights learned from millions of photos, so your classifier starts from real edges and textures instead of random noise.
Why Does This Exist?
Training a ResNet-50 on ImageNet from random weights takes days on 8 GPUs and over a million labeled photos. Almost nobody has that budget for every new task, and almost nobody needs it: the early and middle layers of every trained vision network converge to the same edge, texture and part detectors regardless of the final classes.
Pretrained models package that shared knowledge as downloadable weights. You load a backbone that already sees, attach a small head for your classes, and train on hundreds of images instead of millions. This is what made modern applied vision possible on ordinary hardware, and it is the starting point for transfer learning and fine-tuning.
Think of It Like This
A library card instead of buying every book
Writing a report from scratch by buying and reading every source book takes months and a full shelf. A library card gives you the same knowledge for the cost of the trip: the books were written once, by others, and every reader reuses them.
A pretrained backbone is the library. ImageNet training wrote the books on edges, fur, metal and text. Your task checks out the relevant volumes through a small classifier head instead of rewriting them.
How It Actually Works
What the weights contain
A backbone like ResNet-50 or ViT-B splits into two parts. The backbone holds the visual knowledge: early stages detect edges and color blobs, middle stages detect textures and parts, late stages combine them into object-level patterns. The head maps those patterns to the 1,000 ImageNet classes. You keep the backbone and discard the head, since your classes differ.
How you use one
- Load the backbone with its published weights, for example
IMAGENET1K_V2in torchvision or a checkpoint from a model hub. - Replace the head with a randomly initialized layer sized to your class count.
- Train the head first with the backbone frozen (feature extraction), then optionally unfreeze top layers with a tiny rate.
The standard ImageNet preprocessing travels with the weights: resize to 256, center-crop to 224, normalize with mean [0.485, 0.456, 0.406] and std [0.229, 0.224, 0.225]. Skip that normalization and accuracy drops several points for no reason, because the weights expect centered inputs.
Code
import torchfrom torchvision import models
weights = models.ResNet50_Weights.IMAGENET1K_V2model = models.resnet50(weights=weights)model.fc = torch.nn.Linear(model.fc.in_features, 2)
preprocess = weights.transforms()Watch Out For
Pretraining domain mismatch
Weights learned from natural photos transfer poorly to X-rays, satellite radar or microscopy, whose pixel statistics differ completely. The symptom is a frozen backbone that plateaus below a small from-scratch model. Match the pretraining domain to your images, for example a medical or remote-sensing checkpoint, or budget for deeper fine-tuning.
Wrong preprocessing silently costs points
Each weight release ships with the exact resize, crop and normalization it trained with. Applying your own normalization shifts every activation off the values the weights expect. Always use the checkpoint's own transform pipeline rather than copying numbers from a blog post.
The Quick Version
- Pretrained vision models reuse backbone weights learned from millions of photos.
- You keep the backbone, replace the head, and train on hundreds of images.
- Early layers hold generic vision knowledge; the head holds task classes.
- Always use the checkpoint's own preprocessing, and check the pretraining domain matches yours.