OCR From Pixels to Text
OCR reads text out of photos in two moves: find where the words are, then decode what the letters say, even when fonts and lighting fight back.
Why Does This Exist?
Invoices, street signs, license plates and handwritten forms lock their meaning in pixels, invisible to search and databases. Typing it all by hand does not scale, and classic template matching shatters on the first unfamiliar font. OCR (optical character recognition) exists to convert pictures of text into strings machines can store and query. It splits into text detection, finding the words, and text recognition, reading them.
Think of It Like This
A librarian with sticky notes
A librarian facing a wall of framed posters first sticks a note on every frame that holds writing (detection), then sits down with each marked frame and transcribes it letter by letter (recognition). She never transcribes unmarked wall, and she never sticks notes without reading what is inside. OCR pipelines keep the same division of labor, which is why a blurry detection box guarantees a garbled transcription downstream.
How It Actually Works
Detection networks propose word or line boxes, handling rotation and curvature that axis-aligned object detection boxes cannot fit. Each crop is rectified to a straight strip, then a recognition network (CNN features plus sequence model plus CTC or attention decoding) emits characters left to right. Document OCR adds layout analysis first; scene-text OCR leans on large pre-training to survive wild fonts.
A worked error rate
A recognizer reads "hello" as "hallo": one substitution across five characters, so the character error rate is 1 / 5 = 20 percent. Word-level accuracy calls the whole word wrong, which is why papers report both: character rate shows the model nearly reads, word accuracy shows the user still retypes. Pipelines budget detection and recognition errors separately because a missed box and a misread letter need different fixes.
Watch Out For
Evaluating on clean scans, deploying on phone photos
Scanned benchmarks flatter models with flat lighting and straight lines, while phone captures bring glare, perspective and motion blur. The symptom is a 20-point accuracy cliff on launch day. Fix it by benchmarking on captures from your actual devices and augmenting training with perspective, blur and glare.
The Quick Version
- OCR is detection (where are the words) followed by recognition (what do they say).
- Detection must handle rotation and curvature, not just straight boxes.
- Recognition reads rectified strips with sequence models and CTC or attention decoding.
- Report character error rate beside word accuracy to separate near-misses from failures.
- Phone-photo conditions, not scans, decide whether a system survives deployment.