Getting the words out of a pile of PDFs is the first step of a lot of work: building a research corpus, loading documents into a search tool or an AI assistant, reviewing a disclosure bundle, moving old content to a new website, or simply grepping a year of invoices for one supplier. Copying and pasting from each file is slow and loses track of which text came from where.
The Extract Text tool processes a whole batch and gives you one plain text file per document, with page markers, in a single ZIP. It reads scans and photos too, with OCR. And for the pictures inside the PDFs, Extract Images does the same in bulk. Nothing is uploaded by either.
How to extract text from many files at once
- Open ihatepdf.cv/extract-text.
- Select or drop all the files together — PDFs, and images such as JPG, PNG, WebP, TIFF, GIF or BMP. With more than one file the tool switches to batch mode.
- Turn on OCR mode if any of the PDFs are scans. Leave it off for PDFs with real text; it is much faster.
- Press "Extract N files".
- Download all: one ZIP with a
name_extracted.txtper file.
What the text files look like
Each PDF becomes one UTF-8 text file. The text of every page follows a marker such as
--- Page 3 ---, so you always know where a passage came from, and you can cite or
return to the right page of the original. Images produce a single block of text.
The text comes out in the order it is stored in the PDF, which for ordinary documents is reading order. On complex layouts — two-column papers, sidebars, pages with many text boxes — lines from different columns can interleave. If you need the structure preserved, convert to Word instead (PDF to Word in bulk rebuilds columns, tables and headings).
Scans and photos in the same batch
You do not need to separate digital PDFs from scans and photos:
- Images are always read with OCR, so a photo of a receipt or a screenshot of a document becomes text.
- PDFs are read directly when OCR mode is off, and every page is OCR'd when it is on. Turn it on for a batch that contains scans.
- A PDF with no extractable text while OCR mode is off is flagged with "no selectable text — turn on OCR mode for scans", so you know exactly which files to rerun.
If what you want from your scans is searchable PDFs rather than text files, batch OCR keeps each page as it looks and adds the text on top.
Extracting images in bulk
Extract Images pulls out the photos, logos and figures embedded in each PDF — the original image data, not a screenshot of the page. In a batch:
- JPEGs come out untouched, byte for byte as they were placed in the PDF, at their full original resolution. Other image types are saved losslessly as PNG.
- Each PDF's images go in their own folder in the ZIP, named with the page
they appear on, such as
brochure/brochure-p4-02.jpg. - "Skip images smaller than" leaves out the tiny spacers, rules and bullet graphics that PDFs are full of — 64 pixels is a good default, 200 for photos only.
- Duplicates are merged, so a logo repeated on every page is saved once.
- PDFs without images are skipped rather than failed.
A small number of PDFs store images in JPEG 2000 or JBIG2, two rare formats browsers cannot decode; those images are reported as skipped. More on how extraction works: extracting images from a PDF at full quality.
Choosing between text, Word and searchable PDF
- Plain text when the words are going into another tool — a search index, a spreadsheet, a script, an AI model — and layout does not matter.
- Word documents when someone is going to edit the content and the layout, tables and headings matter.
- Searchable PDFs when the documents stay as documents and just need to be findable.
Working with the output
Plain text is the most portable format there is. The ZIP's files can be searched in one go with
your operating system's search, a code editor's "find in files", or grep on the command
line; imported into a spreadsheet; or loaded into a notes app or a document assistant. Because each
file keeps its document's name and page markers, any line you find leads straight back to its
source.
Frequently asked questions
How do I extract text from multiple PDF files at once?
Open ihatepdf.cv/extract-text, select or drop all the PDFs together and press Extract. Each PDF becomes its own .txt file and you download them all as one ZIP.
Can it extract text from scanned PDFs?
Yes. Turn on OCR mode and every page is read with OCR. Images such as JPG and PNG are always read with OCR.
Does the text keep page numbers?
Yes. Each page's text follows a "--- Page N ---" marker, so you can trace every passage back to its page.
Can I extract all the images from multiple PDFs?
Yes, with Extract Images. Each PDF's embedded images are saved at their original quality, in a folder per PDF, with small decorative images skipped.
Why is the text from a two-column PDF mixed up?
Plain text follows the order the PDF stores it, which on multi-column layouts can interleave lines. Convert to Word to keep columns and structure.
Are my files uploaded?
No. Text and image extraction both run in your browser, and the files never leave your device.