PDFArrow · Scanned PDF field guide

Choose OCR settings by scan quality and document type

OCR quality depends more on the source image and verification process than on a single accuracy percentage.

Reviewed July 21, 2026 by PDFArrow Product Engineering.

Start with the source

Rescan at roughly 300 DPI when possible. Keep pages flat, fill the frame, use even light, remove shadows, and select the main document language. Cleanup can reduce noise but cannot recreate letters that were never captured.

Use confidence as a review signal

A high average confidence can still hide a wrong account number or name. Compare critical fields against the image, especially short strings, punctuation, decimal points, dates, and characters such as O/0 or I/1.

Preserve the original

Treat the searchable PDF and TXT as working derivatives. Keep the scan as the source of truth because OCR adds a machine-generated text layer and may misread layout or characters.

Recommended OCR approach by source type

SourceExpected resultRecommended setupReview focus
Clean printed pageStrong for normal fontsAuto cleanup, correct language, orientation onNames, dates, numbers
Phone photoVariable around glare and perspectiveRetake square to page, even light, crop edgesShadow lines and curved text
Form or tableWords usually stronger than cell orderHigh resolution, minimal cleanupRow/column association and checkboxes
ReceiptVariable for faded thermal printIncrease contrast, keep full widthDecimals, currency, totals, dates
Faint photocopyLower confidenceRescan darker before aggressive cleanupDropped characters and broken words
Multilingual printBest when the main language is selectedProcess separate language groups when possibleNames and mixed-language lines
HandwritingUnreliable in this workflowManual transcription and second-person reviewEvery field

Downloads and evidence

Related tools and resources