PDF OCR — Extract Text from Scanned PDFs

Pull text out of a scanned or image-only PDF using optical character recognition.

Click to upload or drag a scanned PDF here Works best on clean, printed-text scans

How to Run OCR on a Scanned PDF

  1. Click the upload area and choose your scanned or image-based PDF.
  2. The first run downloads the OCR engine's language data (about 11MB) — this is cached, so later runs start immediately.
  3. Optionally enter a page range like "1-3, 5" in the Pages field and click "Re-Run OCR" to scan only those pages — leave it blank to scan the whole document.
  4. Each page is rendered as an image and scanned for text, with a live progress bar tracking recognition as it happens.
  5. Review the extracted text and its confidence score, then copy it or download it as a .txt file once it's done.

Frequently Asked Questions

How is this different from the regular Extract Text from PDF tool?

The regular Extract Text tool reads a PDF's embedded text layer, which only exists if the PDF was generated from a real document. This OCR tool instead renders each page as an image and runs optical character recognition on it, which works on scanned documents and image-only PDFs that have no text layer at all.

Why is this slower than the regular text extraction tool?

Each page has to be rendered as a full image and then run through an OCR recognition pass, which is far more computationally expensive than simply reading an existing text layer — expect noticeably longer processing time, especially for PDFs with many pages.

What OCR engine does this use, and what languages does it support?

It uses Tesseract.js, a WebAssembly port of the Tesseract OCR engine, loaded with the English language pack. Text in other languages or non-Latin scripts won't be recognized accurately.

Will this work well on a low-quality scan?

Accuracy depends heavily on scan quality — skewed pages, low resolution, heavy compression artifacts, or handwriting will all reduce accuracy. Clean, high-resolution scans of printed text give the best results.

Is my file uploaded anywhere?

No — page rendering and OCR recognition both run entirely inside your browser. Your PDF is never uploaded to a server.

Can I run OCR on just some pages instead of the whole PDF?

Yes — use the Pages field to enter specific page numbers and ranges, like "1-3, 5, 8-10", and only those pages are rendered and scanned. Leave it blank to run OCR on every page, and use Re-Run OCR to apply a new page range without re-uploading the file.

Does the tool show progress while OCR is running?

Yes — a progress bar tracks Tesseract.js's real recognition progress for the page currently being scanned, combined with how many pages are left, so you can see how far along the whole job is instead of just guessing from the status text.

Does it show how confident it is in the extracted text?

Yes — Tesseract.js reports a confidence score from 0-100% for each page, and once OCR finishes the tool shows the average confidence across all processed pages, giving a rough sense of how reliable the result is overall.