A scanned document is a photograph of a page. It looks like text, but your computer sees pixels, so you cannot search it, select a sentence or copy a figure out of it. One scan is a nuisance. A filing cabinet's worth — years of invoices, a box of medical records, a department's signed contracts — is an archive nobody can search.
Optical character recognition (OCR) fixes that by reading the words on each page. The OCR tool now takes a whole batch of scanned PDFs and returns each one as a searchable PDF, without uploading a single page.
How to OCR many PDFs at once
- Open ihatepdf.cv/ocr-pdf in a modern browser on a computer or a phone.
- Select or drop all the scanned PDFs together. The tool switches to batch mode as soon as it has more than one file.
- Choose the language, or leave it on Auto-detect so each file's writing system is identified from its first page.
- Leave "Enhance scans" on unless the scans are already very clean.
- Press "OCR N files" and keep the tab open and visible while it works.
- Download all. Each file is saved as
name_searchable.pdf.
What "searchable" means
The output keeps each page as you saw it, with an invisible layer of recognised text placed exactly over the words in the image. Open it in any PDF reader and you can press Ctrl+F (Cmd+F on a Mac) to find a word, select a paragraph and copy it, and your computer's file search can index the document. Nothing about the look of the page changes.
To build that, each page is re-rendered as an image and the text layer is added on top. Pages are rendered at high resolution for small documents and slightly lower for very long ones, so a 300-page book still fits in a browser's memory. The output can therefore be larger or smaller than the original scan depending on how the scan was compressed.
Only run OCR on scans. A PDF that already has real text — anything exported from Word or saved from a web page — gains nothing from it. If you just need the text out of those, use bulk text extraction instead.
Languages, and when to pick one yourself
OCR reads far more accurately when it knows which language to expect. Twenty languages are supported: English, Hindi, French, German, Spanish, Portuguese, Italian, Chinese (Simplified and Traditional), Japanese, Korean, Arabic, Russian, Greek, Hebrew, Thai, Turkish, Polish, Dutch and Swedish.
Auto-detect works per file: it looks at the first page, identifies the writing system — Latin, Devanagari, Arabic, Cyrillic, Chinese, Japanese, Korean, Greek, Hebrew or Thai — and loads the matching model. So a batch that mixes English, Hindi and Arabic documents is handled without sorting. What it cannot do is tell apart languages that share an alphabet: every Latin-script document is read with the English model. That still reads the letters, but it handles accents and language-specific words less well, so for French, German, Spanish, Portuguese, Italian, Polish, Dutch, Swedish or Turkish documents, choose the language yourself. Do the same for Traditional Chinese, since auto-detect assumes Simplified.
Choose a language explicitly, too, when a file's first page is a poor sample — a cover with a logo and two words, or a form that is mostly boxes. If your batch mixes several Latin-alphabet languages, run one batch per language.
What "Enhance scans" does
Enhancement cleans up each page before it is read: it evens out contrast and lifts faint text off grey or yellowed backgrounds. When a page still reads with low confidence, it is tried again in pure black and white, and whichever reading is more confident is kept. That second pass is what rescues faint photocopies and phone photos of documents. On already-clean, high-contrast scans you can switch it off to save a little time.
How long it takes, and keeping it moving
OCR is the heaviest thing a browser does with a PDF: every page is rendered, cleaned up and read word by word. Expect a few seconds per page on a modern laptop and longer on a phone, so a hundred-page batch is a matter of minutes, not seconds. Two things keep it moving:
- Keep the tab visible. Browsers slow down background tabs to save power, and a long OCR run in a hidden tab crawls. Give it its own window and carry on in another.
- Cancel stops at once. If you started the wrong batch, Cancel halts the OCR engine immediately rather than finishing the current file.
Getting better results from scans
- Scan at 300 DPI. Below about 200 DPI small print starts to break apart; above 300 the gain is small and the files are large.
- Straighten sideways pages first. OCR reads upright text; fix scanner orientation with batch rotate before you run it.
- Prefer flat scans to angled photos. A document photographed at an angle can be read, but a flat scan always reads better.
- Expect handwriting to be hit and miss. OCR is built for printed text.
What to do with searchable PDFs
Once a batch has text, the rest of the toolbox opens up. A long scan of many documents can be split wherever a phrase such as "Invoice Number" appears, the text can be pulled into plain files with Extract Text, and the scans can be turned into editable documents with PDF to Word in bulk, which also OCRs scanned pages.
Why a browser, for a batch of scans
The documents people OCR in bulk are rarely trivial: medical records, legal files, signed contracts, tax paperwork. Most online OCR services process them on their servers, and the ones that do batches usually charge for it. Running the OCR engine in your browser means the scans are read on your own device and never sent anywhere — and with no server bill, no per-page charge.
Frequently asked questions
Can I OCR multiple PDFs at once for free?
Yes. Select or drop several scanned PDFs on ihatepdf.cv/ocr-pdf and press OCR. Each becomes a searchable PDF, with no file limit and no sign-up.
Does it detect the language of each file?
It detects each file's writing system from its first page — Latin, Arabic, Cyrillic, Devanagari, Chinese and so on — and loads the matching model. Latin-script files are read as English, so choose French, German, Spanish and other Latin-alphabet languages yourself.
Will the searchable PDFs look different?
No. The page images look the same; an invisible text layer is added on top so you can search, select and copy.
How long does batch OCR take?
Usually a few seconds per page on a laptop, longer on a phone. Keep the tab visible, because browsers slow down background tabs.
Should I OCR PDFs that already have text?
No. OCR only helps scans and image-only PDFs. For PDFs with real text, use Extract Text to get the text out directly.
Can OCR read handwriting?
Only unreliably. It is designed for printed text; neat block capitals sometimes work, cursive rarely does.
Are my scans uploaded?
No. The OCR engine runs in your browser, so every page is read on your own device.