Image Dataloaders
A dataloader turns stored files into steady GPU-ready batches, hiding disk reads, decoding and padding behind parallel workers so the GPU never waits.
Why Does This Exist?
Training reads thousands of JPEG files per minute, decodes them, resizes them, converts them to tensors and stacks them into batches. Done naively in the training loop, the GPU finishes a step in 50 milliseconds and then waits 200 milliseconds for the next batch from disk. Utilization drops under 30% and a job that should take hours takes a day.
The dataloader exists to hide that latency. Worker processes fetch and preprocess the next batches in parallel while the GPU trains on the current one, so steady batches arrive from a queue instead of from disk. This page covers batching, shuffling, workers and memory flow. What goes into each image lives in dataset preparation and per-sample transforms in image augmentation.
Think of It Like This
A kitchen pass feeding a fast chef
A chef plates one dish every minute but each ingredient takes five minutes to fetch and chop. One runner means the chef stands idle. Four runners chop ahead and line finished plates on the pass, so the chef never waits.
The GPU is the chef. Disk reads and JPEG decoding are the fetching and chopping. Workers are the runners and the prefetch queue is the pass.
How It Actually Works
Batching stacks samples of shape into one tensor so the GPU processes them in parallel. A batch of 256 RGB images at 224x224 in float32 holds values, about 147 MB, before gradients.
Three settings control throughput. Workers () set how many processes decode in parallel; 4 to 8 saturates most SSD setups and more workers past that add context-switch cost. Prefetch keeps 2 batches per worker ready so the next batch is already decoded when the step ends. Pinned memory stages batches for fast host-to-device copies.
Shuffling with randomizes order each epoch so batches carry mixed classes. With 50,000 images and batch 256, one epoch is steps; drops the final 80-image remainder so every step has identical shape.
Code
steps = -(-50000 // 256) # ceil divisionprint(steps)# -> 196
remainder = 50000 % 256print(remainder)# -> 80
batch_mb = 256 * 3 * 224 * 224 * 4 / (1024 ** 2)print(round(batch_mb, 1))# -> 147.0Watch Out For
Setting num_workers to zero on a real dataset
The default single-process loading decodes JPEGs in the training loop. The symptom is low GPU utilization with idle gaps in the profiler timeline. Raise workers to 4 or 8 and confirm utilization climbs before touching the model.
Leaking augmentation across workers through a shared random seed
Forked workers can inherit identical random states, so every worker applies the same crop. The symptom is repeated identical images inside one batch. Seed each worker from its id so transforms stay independent.
The Quick Version
- Dataloaders overlap disk reads and decoding across workers so the GPU trains instead of waiting.
- A batch of 256 RGB images at 224 resolution is about 147 MB of float32 input before gradients.
- Set 4 to 8 workers with prefetching and pinned memory, then check the profiler for idle gaps.
- Shuffle every epoch and use drop-last so each of the 196 steps per 50,000 images has identical shape.
- Identical crops inside one batch point at a shared worker seed, not at the model.