PDFArrow · Scanned PDF field guide
Choose OCR settings by scan quality and document type
OCR quality depends more on the source image and verification process than on a single accuracy percentage.
Reviewed July 21, 2026 by PDFArrow Product Engineering.
Start with the source
Rescan at roughly 300 DPI when possible. Keep pages flat, fill the frame, use even light, remove shadows, and select the main document language. Cleanup can reduce noise but cannot recreate letters that were never captured.
Use confidence as a review signal
A high average confidence can still hide a wrong account number or name. Compare critical fields against the image, especially short strings, punctuation, decimal points, dates, and characters such as O/0 or I/1.
Preserve the original
Treat the searchable PDF and TXT as working derivatives. Keep the scan as the source of truth because OCR adds a machine-generated text layer and may misread layout or characters.
Recommended OCR approach by source type
| Source | Expected result | Recommended setup | Review focus |
|---|---|---|---|
| Clean printed page | Strong for normal fonts | Auto cleanup, correct language, orientation on | Names, dates, numbers |
| Phone photo | Variable around glare and perspective | Retake square to page, even light, crop edges | Shadow lines and curved text |
| Form or table | Words usually stronger than cell order | High resolution, minimal cleanup | Row/column association and checkboxes |
| Receipt | Variable for faded thermal print | Increase contrast, keep full width | Decimals, currency, totals, dates |
| Faint photocopy | Lower confidence | Rescan darker before aggressive cleanup | Dropped characters and broken words |
| Multilingual print | Best when the main language is selected | Process separate language groups when possible | Names and mixed-language lines |
| Handwriting | Unreliable in this workflow | Manual transcription and second-person review | Every field |