Optical character recognition

LLM foundationsRetrieval and dataPublished By Simon Budziak

Optical character recognition, or OCR, converts text visible in scanned pages, photographs, or image-only PDFs into machine-readable characters. It makes documents searchable and editable, but it does not by itself understand the document's meaning. Accuracy varies with image quality, handwriting, layout, language, and typography.

What does OCR produce?

OCR detects text regions, recognizes characters, and may return positions and confidence values. The output is a transcription of visible marks, not a verified interpretation. The US digitization guidelines define OCR as converting raster image characters into digitally coded text. That text can feed document chunking and search.

When does OCR need another layer?

Use Document AI when the workflow needs document type, fields, tables, or relationships. Compare critical values with the page image before an automated action, especially for totals, dates, identifiers, and low-confidence spans. A multimodal LLM may interpret visual context, while structured output keeps extracted fields machine-checkable.

Frequently asked questions

Can OCR read handwriting?

Some systems can, but handwriting recognition is usually less reliable than clear printed text and needs representative testing.

Does OCR understand tables and forms?

Basic OCR returns text. Document understanding systems add layout, table, field, and relationship extraction.

Summarize this page with

Train your team to build this