Image to Text (OCR)
Extract text from a screenshot, scan, or photo β the recognition engine runs entirely in your browser.
How to Extract Text from an Image (OCR)
- Click the upload area or drag in a JPG, PNG, or WebP image containing text β screenshots, scans, and photos all work.
- Optionally check "Convert to grayscale" or "Boost contrast" if the image has uneven lighting or low-contrast text β both are applied on-device before OCR runs.
- Click "Extract Text" β the first time you do this, the browser downloads a one-time OCR language pack (about 11MB).
- Watch the progress bar for live recognition progress as the OCR engine analyzes the image.
- Once done, review the extracted text and its confidence score, then click "Copy Text" to copy it to your clipboard.
Frequently Asked Questions
Why does the first extraction take so much longer than later ones?
The tool uses Tesseract.js, a WebAssembly OCR engine that needs to download its English language training data (roughly 11MB) before it can recognize anything. Your browser caches that data after the first download, so every extraction after that starts almost immediately.
What languages does this OCR tool support?
Currently only English - the recognition worker is loaded with the eng language pack specifically. Text in other languages or scripts won't be recognized accurately.
Why did I get garbled or missing text from my image?
OCR accuracy depends heavily on image quality - low resolution, blurry photos, unusual fonts, low contrast, or heavily stylized/handwritten text all reduce accuracy. Screenshots and scans of printed text with good contrast give the most reliable results.
Is my image uploaded to a server to run OCR?
No. Tesseract.js runs the entire recognition process inside your browser using WebAssembly - your image and the extracted text never leave your device.
Can it read handwriting?
Not reliably. Tesseract's OCR model is trained primarily on printed and typed text; handwritten text recognition is a different, much harder problem and results will generally be poor.
Does the tool show how confident it is in the extracted text?
Yes. After each extraction, Tesseract.js reports a confidence score from 0-100% indicating how certain the OCR engine is about the recognized text overall. Lower scores usually mean blurrier images, unusual fonts, or low contrast.
What do the grayscale and contrast boost options do?
They're optional preprocessing filters applied to your image with the HTML canvas before OCR runs. Grayscale removes color information that can confuse the recognizer, and contrast boost widens the gap between text and background. Both can improve accuracy on photos with uneven lighting or low-contrast text - try toggling them if your results look off.