PDF OCR
Recognise the text in a scanned PDF and get back a file you can search and copy from.
Recognition runs on your device, so scanned records and archives stay with you. Only the language data is fetched, once, and cached by your browser.
Convert the searchable result to Word
About PDF OCR
A scanned PDF is a picture of a document. You can read it, but nothing else can: search finds nothing, copy gives you nothing, and every tool that works on text has nothing to work with. Text recognition fixes that by reading the shapes of the letters and writing what it finds back into the file as an invisible layer sitting exactly over the printed words. The page still looks identical — the scan is untouched — but the document is now searchable, selectable and convertible. Recognition runs in your browser, in any of twelve languages, and the extracted text is available on its own as well.
Features
- Recognition runs in your browser — pages are never uploaded
- Twelve languages, including Hindi, Russian, Chinese, Japanese and Arabic
- Invisible text layer sits over the original scan, which is left as it is
- Two scan qualities, trading speed against accuracy on small print
- Plain text output alongside the searchable PDF
- Per-page progress, with a stop control for long documents
- Language data downloads once and is cached by your browser
How to use the PDF OCR
- Drop the scanned PDF onto the page
- Choose the language printed on the page and a scan quality
- Start recognition and watch it work through the pages
- Download the searchable PDF, or just the recognised text
Example
Input
1987-minutes.pdf — 8 scanned pages, typewritten English
Output
1987-minutes-searchable.pdf — same images, now with selectable text behind them.
The visible page does not change; what changes is that the words are now in the file.
Common errors & troubleshooting
- The recognised text is full of mistakes. — Recognition depends on the scan. Try the higher quality setting, and check the language matches the page — English data on a German page produces exactly this.
- Recognition is very slow. — It is genuinely heavy work, and it is happening on your own processor rather than a server. Use the faster quality, and split long documents.
- A document with two languages recognises badly. — One language is used per run. Split the document by language, recognise each part, and merge the results.
- Handwriting was not recognised. — This recognises printed and typewritten text. Handwriting needs a different class of model entirely.
Frequently asked questions
- Does the recognised text replace the scan?
- No. The image stays exactly as it was and the text is added behind it, invisible. That is what makes a searchable PDF look identical to the original.
- Are my scanned pages uploaded for recognition?
- No. The recognition engine runs in your browser. The one thing that is fetched is the language data, once, and your browser caches it afterwards.
- How many pages can I recognise at a time?
- Up to thirty in one run, because recognition is slow enough that longer documents are better handled in batches.
- Which quality setting should I choose?
- Start with the faster one. Move to the higher setting when the print is small, the scan is old, or the first attempt came back messy.
- Can I OCR a PDF that already has text?
- You can, but there is no point — check first by trying to select a word. If text is already there, convert it directly instead.
Related tools
- PDF to Text — Extract selectable text from a PDF as plain text or Markdown.
- Scan to PDF — Use a camera as a document scanner and save the pages as a PDF.
- PDF to Word — Rebuild a PDF as an editable .docx document.
- PDF to Excel — Pull the tables out of a PDF into an .xlsx workbook or CSV.
- PDF to HTML — Convert a PDF into clean, searchable, responsive HTML.
- Compress PDF — Shrink a PDF by re-rendering each page to a JPEG at a chosen quality and resolution.
- PDF Summarizer — Summarise a long PDF section by section, on your device.
All ArrayKit tools