Numbering Scanned PDFs: OCR, Searchability and Image Pages
Why do scanned pages behave differently?
A born-digital PDF already contains text objects, so searching it works immediately. A scan contains a picture of a page, and to a computer it is no more searchable than a photograph. OCR examines that picture and generates a text layer positioned behind the image, which is what makes the words findable.
This distinction drives most of the surprises in a scanned production. A file can look identical on screen to a born-digital one and still return nothing when searched, and a reviewer who assumes otherwise may conclude a document does not mention a term when the text layer simply does not exist.
Is the stamped identifier searchable?
Usually yes, because tools normally add the identifier as a text object drawn over the page rather than as part of the image. That is why a document can be searchable by its Bates number while its actual content is not, which is a common and confusing state for a set of scans that never went through OCR.
Verify rather than assume, because some workflows flatten output to an image, which converts the stamp into pixels and removes its text layer along with everything else. Searching a delivered file for one known identifier is a fast check that catches this before it reaches the other side.
Should OCR run before or after numbering?
Running OCR first is the usual choice, because it means the recognition engine sees only the original page and cannot mistake the production markings for document content. If OCR runs after stamping, the identifier and any confidentiality endorsement may be recognised and added to the text layer, so a search for a number returns matches from pages where it was merely printed in the corner.
Where a workflow forces the reverse order, the practical mitigation is to confirm what ended up in the text layer on a sample of pages, and to keep the stamped identifier as a genuine text object rather than an image so that it can be distinguished.
What affects OCR accuracy on production scans?
Scan resolution, contrast and page condition, more than the software. Faint photocopies, skewed pages, handwriting, stamps overlapping text and unusual typefaces all reduce accuracy, and OCR generally reports no error when it guesses badly, so poor results are silent rather than obvious.
Because accuracy is uneven, protocols often address it directly, specifying whether OCR text must be delivered, in what form, and to what standard. Sampling a set of pages and reading the extracted text against the image is the only dependable way to know what you are actually producing.
| Case | Identifier searchable | Page content searchable |
|---|---|---|
| Born-digital PDF, stamped | Yes | Yes |
| Scan, OCR run, stamped | Yes | Yes, to OCR accuracy |
| Scan, no OCR, stamped | Yes | No |
| Scan, output flattened to image | No | No |
| Scan, OCR run after stamping | Yes | Yes, but may include the stamp |
Common questions
Does stamping a scan make it searchable?
Only for the stamp. The identifier is normally added as text and becomes findable immediately, but the page content stays an image until OCR generates a text layer. A set that is searchable by number and returns nothing for its content is almost always scans that never went through OCR.
Why does searching return pages that only show the number in the corner?
Because OCR ran after stamping and read the production markings as document content. The identifier then appears in the text layer of that page as though it were part of the document. Running OCR before numbering avoids it, and sampling the extracted text confirms which order was used.
Should OCR text be delivered with the production?
That is a protocol question rather than a technical one. Many ESI protocols require extracted or OCR text alongside images, sometimes as page-level files referenced from the load file. Because it affects deliverables and volume, settle it before preparing the production rather than at delivery.
Does flattening the output cause a problem?
It can. Flattening converts everything on the page, including the stamped identifier, into pixels, so the file becomes entirely unsearchable. If a workflow flattens by default, confirm the delivered files still return a match when searched for a known identifier.
How accurate does OCR need to be?
Accurate enough for the use the protocol expects, which is why sampling matters. Recognition quality varies with scan resolution, contrast, skew and handwriting, and failures are silent rather than flagged. Reading extracted text against the page image on a sample is the only reliable measure.