OCR & text

What is OCR? Optical character recognition explained simply

OCR turns images of text into real, searchable text. Learn how it works, when to use it, how accurate it is and how to OCR a PDF or photo for free.

Key takeaways

  • OCR (optical character recognition) converts pictures of letters into real text a computer can search and copy.
  • A scanned PDF without OCR is just images — you can't search it or select words.
  • Accuracy depends mostly on image quality and choosing the correct language.
  • Osd-Scan's OCR supports 24 languages and runs on your device, without uploading files.

You scan a contract, open the PDF and press Ctrl+F to find a clause — nothing. The text is there for your eyes, but not for the computer. That's because a scan is a picture. OCR is the technology that turns that picture into real, searchable text.

OCR in one sentence

Optical character recognition (OCR) analyses an image, finds the shapes of letters and words, and converts them into machine-readable text.

How OCR works

  1. Cleanup. The image is straightened and converted to high-contrast black and white.
  2. Layout analysis. The engine finds text blocks, lines and words, and separates them from pictures and tables.
  3. Recognition. Each word is compared with patterns the engine has learned for the selected language.
  4. Language model. Dictionaries and statistics fix likely mistakes — "rn" vs "m", "0" vs "O".
  5. Output. The text is returned, often with the exact position of each word on the page.

Searchable PDF: the best of both worlds

When you OCR a PDF, the result is usually a searchable PDF: the page image stays exactly the same, and an invisible text layer is placed precisely over the words. You see the original scan; your PDF viewer can search, select and copy the text.

When to use OCR

  • Archiving — find any document later by searching for a name, invoice number or amount.
  • Copying text from a letter or book without retyping it.
  • Editing — convert a scanned PDF to Word, then edit it. See convert a scanned PDF to Word.
  • Accessibility — screen readers can read OCR text aloud to blind and low-vision users.

How accurate is OCR?

FactorEffect
Sharp, well-lit scanHigh accuracy
Correct language selectedEssential — wrong language = many errors
Small or decorative fontsMore mistakes
Blur, shadows, skewMore mistakes
HandwritingUsually poor

How to OCR a document for free

  • New scan: in Scan to PDF, tick Recognize text (OCR) before saving.
  • Existing PDF or photo: open OCR PDF, add the file, choose the language(s) and download the searchable PDF.
  • Just the text: use PDF to Text with OCR enabled to get a .txt file.

All of these run in your browser. The OCR engine is downloaded once and your documents are never uploaded.

OCR and privacy

Cloud OCR services read your documents on their servers. That's a concern for medical letters, contracts or IDs. On-device OCR, like Osd-Scan's, keeps every word on your device.

OCR turns a stack of pictures into a searchable archive. Once you've tried searching your scans, you won't want to save a PDF without it again.

Frequently asked questions

Is OCR 100% accurate?

No. With a clean, sharp scan of printed text, accuracy is usually very high, but small fonts, blur, unusual fonts and handwriting cause errors. Always proofread important numbers.

Does OCR change how my PDF looks?

No. OCR adds an invisible text layer behind the image, so the page looks exactly the same but becomes searchable.

Which languages are supported?

24 languages including English, Hebrew, Arabic, Urdu, Hindi, Persian, Russian, French, German, Spanish, Portuguese, Turkish, Chinese, Japanese and Korean.

Can OCR read handwriting?

Neat printed handwriting sometimes works; cursive usually doesn't. OCR is designed for printed text.