Applies to: Searchable scans; recognition accuracy depends on source quality and language. No OCR engine accuracy benchmark is claimed.

THE SHORT ANSWER

Run OCR in a trusted desktop scanner or PDF application, select the document language, keep the original page image, then search for several known words and verify critical names and numbers.

Why scanned text is not searchable

A basic scan records each page as an image. The letters look readable to a person, but the PDF may contain no actual text characters. OCR analyzes the image and adds recognized text.

Searchable-image output preserves the visual page and places an invisible text layer behind it. This is usually preferable for records because the original appearance remains visible.

Prepare the scan

OCR works best with straight pages, even lighting, clear contrast, and enough resolution. Use about 300 DPI for ordinary printed text. Select the correct language so the software expects the right characters and dictionary patterns.

  1. Rotate every page upright.
  2. Crop dark borders and unrelated background.
  3. Choose the correct language or language combination.
  4. Run OCR across all pages.
  5. Save as a new file.
  6. Test search, copy, and reading order.

Verify important information

OCR can confuse 0 and O, 1 and l, punctuation, accented characters, or text crossing stamps and folds. Search for several words from different pages and copy a paragraph into a text editor to inspect it.

For financial, legal, historical, or identity records, verify critical fields against the visible scan. OCR output is a convenience layer, not proof of accuracy.

Privacy and online OCR tools

Uploading a document to an online service gives that service access to its contents. Review the provider’s privacy and deletion terms before uploading confidential, personal, medical, legal, or business records.

Use an offline application or organization-approved service when the document contains sensitive information.

Score a small field sample rather than trusting search alone

Choose ten short fields whose correct values you can read visually: names, dates, totals and identifiers. Compare each recognized value against its page image and count exact matches. Eight correct fields out of ten is 80% on this small sample only; it does not establish the accuracy of the whole file.

A search hit proves that a word exists in the text layer, not that it is aligned with the right visible page or that reading order is sensible. Copy a two-column passage and inspect the order. Change recognition language or recapture a blurred page before repeatedly running OCR over the same bad image.

Sources and further reading

Vendor instructions and background references are linked below. Worked examples are our explanatory calculations, not vendor performance claims. Check the instructions for your current software and device.

Report a correction with this page’s title and the step or result that needs attention.

This guide provides general document-workflow information. Keep an unchanged original before resizing, compressing, or converting an important file.