Text Recognition From Crops
Text recognition reads a cropped word strip left to right, turning pixel columns into characters without needing anyone to cut the letters apart.
Why Does This Exist?
Cutting a word into individual letters before reading fails on cursive, touching print and stylized fonts where letter boundaries do not exist. Yet text detection only delivers whole-word strips. Text recognition exists to read unsegmented strips end to end: a CNN reads visual columns, a sequence model adds context, and a decoder emits characters. It is the second half of OCR.
Think of It Like This
Reading a train timetable through a slit
Slide a slit window across a timetable row and call out what passes: digits, dashes, blanks. You never cut the row apart; position plus memory of what came before disambiguates a smudged 3 from an 8. CTC decoding works like the caller, collapsing repeats and blanks into the final string, while attention decoding glances back at the informative columns. The analogy stops at training: the caller learns from rules, the network from thousands of strip-and-transcript pairs.
How It Actually Works
The strip is rectified to a fixed height, a CNN emits a left-to-right feature sequence, and optional recurrent or transformer layers mix context across positions. CTC decoding aligns frame predictions to the transcript by allowing blanks and repeats, training without character positions. Attention decoders instead emit one character at a time while weighting the relevant columns, which handles irregular layouts at higher compute cost.
A worked CTC collapse
Frame predictions for a strip read h, h, blank, a, a, l, l, o. CTC first merges consecutive repeats within the blank-separated runs to h, a, l, o, then drops the blank, yielding "halo". Fourteen raw frame outputs compress into four characters with no position labels supplied. The same collapse turns doubled predictions from wide letters into single characters instead of stutters.
Watch Out For
A vocabulary that never saw your characters
Models trained on alphanumeric English emit garbage on accents, currency symbols and rare glyphs because those classes never existed in training. The symptom is confident substitution of e for é across whole documents. Fix it by extending the charset and fine-tuning on samples that contain your real symbols.
The Quick Version
- Recognition reads whole word strips, never pre-cut letters.
- CNN features plus sequence context feed CTC or attention decoders.
- CTC trains without character positions by collapsing repeats and blanks.
- Attention decoding handles irregular text at higher cost.
- The training charset must contain every symbol your deployment will meet.