How to Make a Scanned PDF Searchable with OCR
A scanned PDF is a stack of pictures: you cannot search it, copy from it, or have it read aloud. OCR adds an invisible text layer behind each page, so the pages look the same but work like a text PDF.
What OCR changes, and what it leaves alone
Optical character recognition finds the letters in each page image and places matching invisible text in the same positions. The original scan is kept, so the document looks exactly as before. Ctrl+F, copy and paste, and screen readers now work, and PDF to Word or PDF to Excel can read it too.
Get a scan OCR can read
- Scan at 300 DPI; lower resolutions lose small print.
- Keep pages straight and flat; heavy skew and curled book pages reduce accuracy.
- Black-and-white or greyscale is fine and gives smaller files.
- Choose the document's language, because recognition uses a dictionary for each language.
Check the result
Search for a few words from different pages, including numbers and names. Copy a paragraph and compare it with the page. ToolkitPoint's OCR PDF runs Tesseract in your browser, so confidential scans are not uploaded. The trade-off is that a long document takes a while on a slow computer.
Limits
Handwriting is not recognised reliably. Stamps, low-contrast photocopies, and decorative fonts cause errors. The text layer uses standard fonts, so some accented or non-Latin characters may be stored as a placeholder. If a page already has real text, it does not need OCR.
Primary sources and further reading
Last reviewed: September 28, 2026