What is OCR PDF?
OCR (optical character recognition) detects text in images of pages and turns it into real, selectable and searchable text.
How to use OCR PDF
- 1
Choose a scanned PDF or an image.
- 2
Select the document's language(s) and a quality level.
- 3
Click Run OCR and download the searchable PDF and a plain-text file.
Features
- Tesseract OCR running on your device
- Multi-language recognition (up to three at once)
- Keeps the original page appearance
- Exports a searchable PDF and a .txt file
Supported formats
- JPG
- PNG
Maximum file size: 100 MB per file.
Practical uses
- Search inside scanned contracts and invoices.
- Copy text from a photographed document.
- Prepare scans for summaries, translation or PDF to Word.
Limitations
- Accuracy depends on scan quality; handwriting is generally not recognized.
- The first run downloads the language model (a few MB), which is then cached.
- Complex layouts such as tables may be recognized in reading order only.
Frequently asked questions
Is OCR 100% accurate?
No OCR is perfect. Clean, straight scans at 200–300 DPI give the best results; check important numbers manually.
Are my scans uploaded?
No. Recognition runs in your browser. Only the language model file is downloaded from a public CDN.
Which languages are supported?
Spanish, English, Portuguese, French, German, Indonesian, Italian and Dutch.
