Skip to content
AI360Xpert
Beta

Fine-Tuning Vision Models

Once a backbone already knows edges and textures from ImageNet, you retrain its top layers on your own small dataset with a tiny learning rate instead of starting from random weights.

Fine-tuning unfreezes the top layers of a pretrained backbone and retrains them gently, adapting ImageNet features to a new task.
Fine-tuning unfreezes the top layers of a pretrained backbone and retrains them gently, adapting ImageNet features to a new task.

Why Does This Exist?

A ResNet-50 trained on ImageNet has seen over a million photos and already knows edges, textures and object parts. Your dataset of 800 factory-defect photos cannot teach that from scratch: a deep network trained on 800 images memorizes them within a few passes and fails on the next batch off the line.

Transfer learning gives you the starting weights. Fine-tuning is the second half of that deal: after training a new classifier head on frozen features, you unfreeze the top backbone layers and keep training with a learning rate around 10 to 100 times smaller than normal. The head learns your classes, then the backbone bends its high-level features slightly toward your domain.

Think of It Like This

Tailoring a suit instead of sewing one

A rack suit already knows what shoulders and lapels are. A tailor takes it in at the waist and shortens the sleeves, small changes near the surface, and hands it back the same day. Nobody unpicks the shoulder seam, the most expensively right part, unless the fit truly demands it.

Fine-tuning works the same way. The frozen backbone is the rack suit. Training the head is the waist adjustment. Unfreezing the top layers with a tiny learning rate is letting out the cuffs a little. Reaching for a big learning rate is unpicking the shoulders: you destroy the part that was already right.

How It Actually Works

The two phases

Phase 1, linear probe. Freeze every backbone weight and train only the new head. With a ResNet-50 backbone that head is one linear layer from 2048 features to your class count, a few thousand weights against your few hundred images. It converges fast and barely overfits because there is almost nothing to fit.

Phase 2, fine-tuning. Unfreeze the top block (layer4 in a ResNet) and train the head plus those layers together with a learning rate near 1e-5, while the head may keep 1e-3. Early layers stay frozen: they hold generic edges that your data cannot improve.

A concrete schedule

Say you have 800 labeled images in 2 classes. Train the head alone for 5 to 10 epochs until validation accuracy plateaus near, say, 88%. Then unfreeze layer4, drop the backbone learning rate to 1e-5, and train 5 more epochs. A typical result is 88% rising to 91 to 93%, because the top features reshape around your defect textures instead of staying whatever ImageNet left behind. Unfreezing the whole network on 800 images usually moves validation the other way.

Code

import torchfrom torchvision import models
model = models.resnet50(weights=models.ResNet50_Weights.IMAGENET1K_V2)model.fc = torch.nn.Linear(model.fc.in_features, 2)
# Phase 1: freeze the backbone, train the head onlyfor name, param in model.named_parameters():    param.requires_grad = name.startswith("fc.")
head_opt = torch.optim.Adam(model.fc.parameters(), lr=1e-3)
# Phase 2: unfreeze the top block, tiny learning rate for old weightsfor name, param in model.named_parameters():    if name.startswith("layer4."):        param.requires_grad = True
full_opt = torch.optim.Adam([    {"params": model.fc.parameters(), "lr": 1e-3},    {"params": model.layer4.parameters(), "lr": 1e-5},])

Watch Out For

A normal learning rate erases the pretraining

Unfreezing with the same 1e-3 rate you used for the head rewrites the backbone in the first epoch. Validation accuracy jumps, then collapses below the linear-probe baseline: classic catastrophic forgetting. The fix is mechanical. Keep backbone rates at 1e-5 or below, and confirm the frozen-head baseline first so you can tell forgetting apart from a bad head.

Unfreezing everything on a tiny dataset

With 800 images and 23 million unfrozen weights, the network has roughly 29,000 weights per image and memorizes. The symptom is a widening train-validation gap right after unfreezing. Unfreeze one block at a time from the top, and stop when validation stops improving.

The Quick Version

  • Fine-tuning unfreezes the top layers of a pretrained backbone and retrains them with a learning rate near 1e-5.
  • Phase 1 trains only the new head on frozen features; phase 2 adapts the top block to your domain.
  • Early layers stay frozen because your small dataset cannot improve generic edge detectors.
  • Too large a learning rate causes catastrophic forgetting; unfreezing too much causes memorization.