CRNN Text Reading Networks
CRNN reads word photos with a CNN eye and an RNN memory, so letters it cannot quite see are rescued by the letters around them.
Why Does This Exist?
Reading "rn" versus "m" in a blurry plate needs neighbours: the word's context disambiguates what pixels cannot. Pure CNN classifiers read each column alone and fail exactly there. CRNN (Convolutional Recurrent Neural Network) exists to add memory to the eye: convolutional layers scan the strip into a feature sequence, bidirectional LSTMs mix left and right context, and CTC decoding emits the string without letter positions. It is the workhorse behind text recognition inside OCR.
Think of It Like This
A proofreader sounding out a smudge
A proofreader covering all but one letter of a smudged word reads it by sounding the neighbours: "_ight" with an l before it is "light". The CNN sees the smudge; the recurrent layers supply the sentence memory; CTC writes the verdict. The analogy stops at the direction: the proofreader reads both ways at once, while the bidirectional LSTM literally does, carrying past and future context into every position.
How It Actually Works
The rectified strip passes through stacked convolutions and a map-to-sequence layer that slices feature maps into left-to-right column vectors. Two-layer bidirectional LSTMs contextualize each column with its neighbours, and a per-frame softmax over the alphabet plus blank feeds CTC, which trains on strip-transcript pairs with no alignment labels and collapses repeats at inference.
A worked rescue
A strip of "night" smudges the g into a blob the CNN scores as q 0.4, g 0.35, blank 0.25. Neighbouring columns confidently read n-i-h-t, and the bidirectional context shifts the blob frame to g 0.55, q 0.25. CTC then collapses the frame sequence n, i, g, h, t into "night". Context moved 0.2 of probability mass, which is the entire margin between a misread and a correct word.
Watch Out For
Fixed-height strips that squash tall scripts
Resizing every strip to one height mangles ascenders, descenders and stacked scripts. The symptom is fine Latin accuracy with collapsing performance on mixed-case and tall fonts. Fix it by preserving aspect ratio with padding and by training on the script heights you will actually serve.
The Quick Version
- CNN layers scan; bidirectional recurrence contextualizes; CTC decodes.
- No character positions needed for training, only strip transcripts.
- Context rescues ambiguous letters that isolated columns misread.
- The design is light enough for on-device plates and receipts.
- Curved and heavily distorted text needs rectification or attention successors.