No. 15
OCR — Image & PDF to Text
Extract text from a photo, scan, or screenshot in English, Malayalam, Hindi, Tamil, and more — processed on your device.
Uploaded 0 bytes Loads a language model (~15–30MB per language) once, then works offline for that language. (network monitor unsupported in this browser)
How your data is handled: everything above runs inside this browser tab; your file is never uploaded. Loads a language model (~15–30MB per language) once, then works offline for that language.
Extracting text from a scanned document or photo is exactly the kind of task where uploading the file feels wrong — IDs, bank statements, handwritten notes. This runs Tesseract locally via WebAssembly, so the image and its text never leave the browser.
How to use it
- Drop an image, screenshot, or scanned PDF page.
- Pick the document's language (or languages, if mixed) from the list.
- Click Extract — a progress readout tracks recognition.
- Copy the text, or download it as a .txt file.
Tips & edge cases
- Straight, well-lit, high-contrast scans recognize far more accurately than an angled phone photo — a quick crop and straighten before running OCR pays off.
- Selecting the correct language matters more than image quality in many cases; recognition on the wrong language model produces garbled output even from a clean scan.
- For a scanned PDF that's already partly text (a hybrid scan), check whether text is already selectable before running OCR — you may not need it.
FAQ
- Is my document uploaded to extract the text?
- No — recognition runs locally using Tesseract compiled to WebAssembly. Nothing is sent to a server.
- Which languages are supported?
- English plus Malayalam, Hindi, Tamil, and other Tesseract-supported languages, each loaded as a separate model on demand.
- Can it read handwriting?
- Tesseract is built for printed and typed text; handwriting recognition is unreliable and not the intended use case.