You scanned a contract, opened the PDF, and can't select a single word. That's because a scan is a photograph, not text. OCR is the technology that bridges the gap — here's how it works and when to use it.
Why scanned PDFs have no text
A PDF can store two very different things that look identical on screen: real text (characters the computer knows about) and an image of text (pixels arranged to look like characters). A document exported from Word contains real text — you can select, search and copy it. A scan or a phone photo is just an image; to the computer it's no different from a picture of a cat.
OCR — optical character recognition — analyses that image, finds the shapes that look like letters, and converts them back into real, editable text.
How OCR actually works
- Clean-up. The engine straightens the page, boosts contrast, and removes noise. A crooked, dim photo is the number-one cause of bad results.
- Layout analysis. It detects columns, paragraphs, tables and headings, so words come out in reading order.
- Character recognition. Each glyph is matched against learned letter shapes — modern engines use neural networks trained on millions of fonts.
- Language modelling. Context fixes ambiguity: "c1ick" becomes "click" because the engine knows English words.
Getting the best results
- Scan at 300 DPI — the sweet spot for accuracy. Below 200 DPI, letters lose the detail OCR needs.
- Keep the page flat and straight. Skewed or curved text drops accuracy sharply.
- Good lighting, no shadows when photographing with a phone.
- Prefer clean fonts. Handwriting and decorative fonts remain the hardest cases — expect to proofread.
What to do with a scanned PDF today
If your PDF already contains real text, you don't need OCR at all — our PDF to Text and PDF to Word tools extract it instantly in your browser. If the extraction comes back empty, your file is a scan: run it through an OCR engine first, then extract. Full OCR needs heavy processing that browsers handle poorly, which is why we label it honestly instead of pretending — see how our local-first tools work.
OCR accuracy: what to expect
On a clean 300 DPI scan of printed text, modern OCR reaches 98–99% accuracy — a typo or two per page. On phone photos it drops with blur and shadow, and on handwriting it varies wildly. Always proofread anything legal or financial after OCR; the last 1% of errors loves to hide in numbers and names.