ImagePDF.Tools
PDF Tool

Make scanned PDFs searchable.

Run Tesseract OCR on any scanned PDF. Adds a text layer so you can select, copy, and search the content. Entirely in your browser.

OCR PDF
Processing securely on your device, no data sent to any server

Drop your scanned PDF here

PDF files only, processed entirely in your browser

Privacy Note: We use your browser's hardware to process this file. Your PDF stays on your computer throughout the entire process, nothing is transmitted.
No upload·100% private·Instant

Zero upload · 300 DPI render · Tesseract OCR · Searchable PDF output

How it works

Three steps. Fully searchable.

Upload your scanned PDF

Drop any scanned PDF: receipts, contracts, books, or forms. Works with any scan quality.

OCR runs in your browser

Tesseract OCR reads every page, detecting words, their positions, and confidence scores. No upload required.

Save your searchable PDF

Save a PDF with an invisible text layer. Looks like the original, but text is now selectable and searchable.

Use cases

Who needs searchable PDFs?

OCR unlocks the content trapped inside scanned documents, making them useful for search, accessibility, and further processing.

Scanned contracts and agreements

Make signed contracts searchable so you can find specific clauses, dates, and names with Ctrl+F instead of reading every page.

Archiving physical documents

Digitised paper records become far more useful with OCR. Full-text search across your entire archive in seconds.

Accessibility for screen readers

Screen readers need actual text, not images. OCR creates a proper text layer so visually impaired users can read the document.

Scanned receipts and invoices

Extract amounts, dates, and vendor names from scanned receipts. Searchable invoices make expense reporting much faster.

Books and academic papers

Scanned textbooks and research papers become truly useful once OCR lets you search for terms and copy quotes directly.

Multi-language document processing

Support for 12 languages means OCR works accurately on documents in Spanish, French, German, Japanese, Arabic, and more.

How OCR works under the hood.

Each page of your PDF is rendered as an image using PDF.js (Mozilla's open-source renderer). That image is then passed to Tesseract.js, which runs the OCR engine entirely in a browser Worker thread, detecting every word and its position on the page.

The recognised text is embedded as an invisible layer using pdf-lib, placed precisely over the original scan. The result looks identical to the input but the text is now fully selectable and searchable. Nothing is sent to any server at any point.

Common questions

Questions answered.

What is OCR?
OCR (Optical Character Recognition) reads text from images and scanned documents. It converts image-based text into actual, selectable characters that can be searched, copied, and edited.
Is my PDF uploaded to a server?
No. The entire process runs locally in your browser using Tesseract.js, the industry-standard open-source OCR engine originally developed at HP Labs. Your file never leaves your device.
What is a searchable PDF?
A searchable PDF has an invisible text layer placed over the original scanned images. It looks exactly like the original scan, but the text can now be selected, copied, and found with Ctrl+F.
What languages are supported?
English, Spanish, French, German, Italian, Portuguese, Russian, Chinese (Simplified), Japanese, Arabic, Hindi, and Korean.
Can I convert the result to Word?
Yes. Click Convert to Word after OCR completes. The searchable PDF is passed directly to the PDF to Word tool, no re-upload needed.
How accurate is the OCR?
Accuracy depends on scan quality. Clean, high-resolution scans of typed text achieve 95% or higher accuracy. Handwriting, unusual fonts, or low-resolution scans give lower results.
How long does OCR take?
A typical 10-page scanned document takes 20 to 60 seconds depending on your device speed. Tesseract loads its language model once and then processes each page sequentially.
Does OCR work on handwritten text?
Tesseract is optimised for printed text. Handwriting recognition is partially supported but accuracy varies significantly based on the clarity and style of the handwriting.
You're offline, cached tools still work