Skip to content
AI360Xpert
Beta

Dataset Preparation for Vision

Dataset preparation turns a folder of raw photos into clean, split, sized and normalized tensors, so training sees consistent inputs and the test score means something.

Preparation turns 50,000 raw photos into stratified 40,000/5,000/5,000 splits with one size and normalization, so the test score measures generalization.
Preparation turns 50,000 raw photos into stratified 40,000/5,000/5,000 splits with one size and normalization, so the test score measures generalization.

Why Does This Exist?

A raw image folder is not a dataset. Sizes vary from 200 pixels to 4000, some files are corrupt, a few labels are wrong, near-duplicates sit in both train and test, and one class outnumbers another ten to one. Train on that directly and two things break: the collate step crashes on mismatched shapes, and the test score measures memorized duplicates instead of generalization.

Dataset preparation is the pass that fixes this before augmentation ever runs. It covers cleaning (deduplication, corrupt-file removal, label checks), splitting (stratified train, validation and test with no leakage), sizing (one canonical resolution) and normalization (channel means and standard deviations). This page covers that pipeline. Labeling workflows live under imbalanced data neighbours and per-image transforms live in image augmentation.

Think of It Like This

Prepping ingredients before cooking

A cook lays out washed, chopped and measured ingredients before the stove goes on. Unwashed vegetables carry grit, unmeasured spices ruin the balance, and tasting from the serving plate contaminates the result.

Raw photos are the unwashed produce. Cleaning removes the grit, the split keeps the tasting spoon out of the serving plate, and fixed sizing plus normalization is measuring everything into the same bowls so each batch cooks evenly.

How It Actually Works

1. Clean first

Drop unreadable files, exact duplicates (file hash) and near-duplicates (perceptual hash) so no near-copy lands in both train and test. Spot-check at least 200 random labels by hand; a 5% label error rate on 50,000 images means 2,500 wrong answers the model will faithfully learn.

2. Split with stratification and no leakage

Split by acquisition group (patient, camera, video clip), not by file, then stratify classes inside each group. A worked split of 50,000 images at 80/10/10 gives 40,000 train, 5,000 validation and 5,000 test. With a 2% rare class, that is 1,000 / 100 / 100 rare images per split; a plain random split can leave validation with 60 and test with 140, so stratify to hold the ratio.

3. Fix size and normalization once

Pick one training resolution (224x224 is the common default) and one normalization (ImageNet means [0.485, 0.456, 0.406] with stds [0.229, 0.224, 0.225] when fine-tuning). Store the raw files plus a manifest of splits, sizes and hashes so the dataloader reproduces the same tensors every run.

Code

from collections import Counterfrom sklearn.model_selection import train_test_split
# 50,000 image ids with a 2% rare class: 49,000 common, 1,000 rarelabels = [0] * 49000 + [1] * 1000
train, temp = train_test_split(labels, test_size=0.2, stratify=labels, random_state=0)val, test = train_test_split(temp, test_size=0.5, stratify=temp, random_state=0)
print(len(train), len(val), len(test))# -> (40000, 5000, 5000)print(Counter(val))# -> Counter({0: 4900, 1: 100})

Watch Out For

Splitting files instead of groups

Splitting video frames or patient scans file by file puts near-identical frames in both train and test. The test score looks excellent and the deployed model fails. The symptom is a large gap between offline accuracy and live accuracy. Split by video clip, patient or camera first, then stratify inside.

Normalizing with statistics computed on the test set

Computing channel means over the full folder including test leaks test color balance into training. The symptom is subtle and the fix is strict: compute means and stds on train only, freeze them, and apply the same values to validation and test.

The Quick Version

  • Preparation turns raw photos into clean, split, sized and normalized tensors before augmentation runs.
  • Deduplicate by hash and spot-check labels, because a 5% error rate teaches 2,500 wrong answers on 50,000 images.
  • Split 80/10/10 by acquisition group with stratification so a 2% rare class keeps its ratio in every split.
  • Fix one resolution and train-only normalization statistics, and store the split manifest for reproducibility.
  • Leakage between train and test is the failure that makes every other metric meaningless.