Skip to content
Bates PDFA BATESFAST FIELD GUIDE
The workflowField notesQuestions
Open BatesFast
Bates PDF / FIELD NOTES
Workflow

Numbering Scanned PDFs: OCR, Searchability and Image Pages

Direct answer A scanned page is an image, so it carries no text layer until OCR creates one, and that changes how a stamped identifier behaves. The stamp itself is normally added as text and stays searchable regardless, but the underlying document content will not be findable unless OCR has been run. Order matters less than verification: what a production needs is that the identifier is searchable, that the page text is searchable where the protocol requires it, and that OCR has not been allowed to read the stamped identifier back as if it were document content. Because OCR accuracy varies sharply with scan quality, teams generally run it before numbering, then confirm on a sample that both the identifier and the body text can be found in the delivered files. Sampling a handful of delivered pages and reading the extracted text against the image is the only dependable check, because recognition failures are silent rather than reported.

Why do scanned pages behave differently?

A born-digital PDF already contains text objects, so searching it works immediately. A scan contains a picture of a page, and to a computer it is no more searchable than a photograph. OCR examines that picture and generates a text layer positioned behind the image, which is what makes the words findable.

This distinction drives most of the surprises in a scanned production. A file can look identical on screen to a born-digital one and still return nothing when searched, and a reviewer who assumes otherwise may conclude a document does not mention a term when the text layer simply does not exist.

Is the stamped identifier searchable?

Usually yes, because tools normally add the identifier as a text object drawn over the page rather than as part of the image. That is why a document can be searchable by its Bates number while its actual content is not, which is a common and confusing state for a set of scans that never went through OCR.

Verify rather than assume, because some workflows flatten output to an image, which converts the stamp into pixels and removes its text layer along with everything else. Searching a delivered file for one known identifier is a fast check that catches this before it reaches the other side.

Should OCR run before or after numbering?

Running OCR first is the usual choice, because it means the recognition engine sees only the original page and cannot mistake the production markings for document content. If OCR runs after stamping, the identifier and any confidentiality endorsement may be recognised and added to the text layer, so a search for a number returns matches from pages where it was merely printed in the corner.

Where a workflow forces the reverse order, the practical mitigation is to confirm what ended up in the text layer on a sample of pages, and to keep the stamped identifier as a genuine text object rather than an image so that it can be distinguished.

What affects OCR accuracy on production scans?

Scan resolution, contrast and page condition, more than the software. Faint photocopies, skewed pages, handwriting, stamps overlapping text and unusual typefaces all reduce accuracy, and OCR generally reports no error when it guesses badly, so poor results are silent rather than obvious.

Because accuracy is uneven, protocols often address it directly, specifying whether OCR text must be delivered, in what form, and to what standard. Sampling a set of pages and reading the extracted text against the image is the only dependable way to know what you are actually producing.

What is searchable in each case
CaseIdentifier searchablePage content searchable
Born-digital PDF, stampedYesYes
Scan, OCR run, stampedYesYes, to OCR accuracy
Scan, no OCR, stampedYesNo
Scan, output flattened to imageNoNo
Scan, OCR run after stampingYesYes, but may include the stamp

Common questions

Does stamping a scan make it searchable?

Only for the stamp. The identifier is normally added as text and becomes findable immediately, but the page content stays an image until OCR generates a text layer. A set that is searchable by number and returns nothing for its content is almost always scans that never went through OCR.

Why does searching return pages that only show the number in the corner?

Because OCR ran after stamping and read the production markings as document content. The identifier then appears in the text layer of that page as though it were part of the document. Running OCR before numbering avoids it, and sampling the extracted text confirms which order was used.

Should OCR text be delivered with the production?

That is a protocol question rather than a technical one. Many ESI protocols require extracted or OCR text alongside images, sometimes as page-level files referenced from the load file. Because it affects deliverables and volume, settle it before preparing the production rather than at delivery.

Does flattening the output cause a problem?

It can. Flattening converts everything on the page, including the stamped identifier, into pixels, so the file becomes entirely unsearchable. If a workflow flattens by default, confirm the delivered files still return a match when searched for a known identifier.

How accurate does OCR need to be?

Accurate enough for the use the protocol expects, which is why sampling matters. Recognition quality varies with scan resolution, contrast, skew and handwriting, and failures are silent rather than flagged. Reading extracted text against the page image on a sample is the only reliable measure.

BEFORE YOU CHOOSE

Does BatesFast fit this job?

Number a mixed PDF set while maintaining one sequence across files. Prepare scans and searchable text before final stamping when required.

What to verify

Check rotation, document boundaries, stamp overlap and exported ranges. Confirm OCR and load-file requirements separately.

Prepare scanned and mixed-size files →

Account and export terms

Google sign-in is required. One complete production is free, with downloads available for seven days. Further productions require a US$299 lifetime license per device, excluding tax.

View current BatesFast plans →

This is a product-operated field guide in the PDFImpose network. Confirm the exported result and current terms before paying.