No. 15

OCR — Image & PDF to Text

Extract text from a photo, scan, or screenshot in English, Malayalam, Hindi, Tamil, and more — processed on your device.

Uploaded 0 bytes Loads a language model (~15–30MB per language) once, then works offline for that language. (network monitor unsupported in this browser)

How your data is handled: everything above runs inside this browser tab; your file is never uploaded. Loads a language model (~15–30MB per language) once, then works offline for that language.

Extracting text from a scanned document or photo is exactly the kind of task where uploading the file feels wrong — IDs, bank statements, handwritten notes. This runs Tesseract locally via WebAssembly, so the image and its text never leave the browser.

How to use it

  1. Drop an image, screenshot, or scanned PDF page.
  2. Pick the document's language (or languages, if mixed) from the list.
  3. Click Extract — a progress readout tracks recognition.
  4. Copy the text, or download it as a .txt file.

Tips & edge cases

  • Straight, well-lit, high-contrast scans recognize far more accurately than an angled phone photo — a quick crop and straighten before running OCR pays off.
  • Selecting the correct language matters more than image quality in many cases; recognition on the wrong language model produces garbled output even from a clean scan.
  • For a scanned PDF that's already partly text (a hybrid scan), check whether text is already selectable before running OCR — you may not need it.

FAQ

Is my document uploaded to extract the text?
No — recognition runs locally using Tesseract compiled to WebAssembly. Nothing is sent to a server.
Which languages are supported?
English plus Malayalam, Hindi, Tamil, and other Tesseract-supported languages, each loaded as a separate model on demand.
Can it read handwriting?
Tesseract is built for printed and typed text; handwriting recognition is unreliable and not the intended use case.