How does OCR work? From Tesseract to document AI: turning scanned PDFs and invoices into searchable, structured text

OCR turns images of text into text. How the pipeline works, where Tesseract still fits, how vision-language models changed document AI, how to OCR a PDF for free, how to extract invoice fields you can trust, and how I built a page-batched OCR and speech workspace on Mistral.

Requirements gathering and user flows, on stage
Requirements gathering and user flows, on stage

Key takeaways

  • OCR has three stages: find the text regions, recognise the characters, and reconstruct the reading order and layout. Most failures are in the first and third, not the middle.
  • Tesseract is free and fine for clean, printed, single-column pages. For scans, tables, handwriting and forms, modern vision-language models are far better.
  • Document AI is OCR plus understanding: extracting the fields you need into a schema you defined, validated, with confidence per field.
  • Process large PDFs in page batches with concurrency, keep page numbers with every chunk, and make citations point to pages.
  • Deterministic validation after extraction, totals that add up, dates that parse, IDs that match a pattern, is what makes the output trustworthy.

What is OCR and how does it work?

OCR, optical character recognition, converts images of text, a scanned page, a photo of a receipt, a PDF that is really a picture, into machine-readable text. It works in three stages: detection finds where text is on the page, recognition reads the characters in each region, and layout analysis puts the pieces back in reading order with their structure, paragraphs, columns, tables. Modern systems do all three with neural networks; older ones like classic Tesseract use hand-built image processing for the first and third.

Also asked as: what is ocr · what is ocr technology · how does ocr work · what is ocr in computer · what is ocr in pdf · ocr meaning · optical character recognition explained · how ocr works step by step · what is ocr scanning

The distinction that saves the most time: a "PDF" may already contain text, in which case you do not need OCR at all, only extraction. Check with a PDF library first; PyMuPDF will tell you in one call whether a page has a text layer [11]. OCR is for pages that are only pixels.

I built two document systems in 2026 that sit on top of OCR: DocuLens AI, which extracts and compliance-checks invoices and contracts with schema validation [13], and the OCR & Speech Workspace, which OCRs large PDFs concurrently in page batches and turns each document into a searchable knowledge base with page-level citations [12]. This page is the pipeline underneath both.

Nobody wants OCR. They want the total from the invoice, the clause from the contract, the answer from the manual, with the page number. OCR is the first ten percent of that. Pranjul Rathour, from building DocuLens AI

What are the stages of an OCR pipeline?

Preprocess the image: deskew, denoise, binarise, fix resolution. Detect text regions and their orientation. Recognise the characters in each region into strings with confidence scores. Reconstruct layout: reading order, columns, paragraphs, tables, headers and footers. Post-process: spell-check against a lexicon, normalise numbers and dates, attach page and position to every span. Then either store the text for search or pass it to an extraction step.

Also asked as: ocr pipeline · ocr process steps · ocr preprocessing · text detection vs text recognition · ocr architecture · how to build an ocr system · ocr post processing · document layout analysis

Detection is where CRAFT-style networks replaced connected-component heuristics, finding text at odd angles and sizes [3]. Recognition moved from character classifiers to sequence models, and TrOCR showed a pure transformer encoder-decoder could read lines end to end with pre-trained vision and language models [4]. Layout is where the hardest problems remain, and where document-understanding models entered.

What is Tesseract, and is it still good?

Tesseract is the open-source OCR engine originally developed at HP and later at Google, now maintained as a community project [1]. Version 4 added an LSTM-based recogniser, which made it good on clean, printed, well-scanned text in dozens of languages, and it remains free, offline and scriptable, with OCRmyPDF wrapping it to add text layers to scanned PDFs [10]. It is still the right tool for clean printed pages at zero cost. It is the wrong tool for photos, handwriting, complex tables and forms.

Also asked as: what is tesseract ocr · is tesseract still good · tesseract vs easyocr · tesseract vs paddleocr · best free ocr · open source ocr · tesseract accuracy · how to use tesseract ocr python · pytesseract

Smith's 2007 overview of the engine's architecture is still the best short explanation of how a classic OCR engine thinks about a page [2].

How do vision-language models change OCR?

Vision-language models read the page image directly and output text, structure or answers, without a separate detection and recognition stage. They handle mixed layouts, tables, handwriting and low-quality scans far better than classic engines, they can output Markdown with headings and tables preserved, and they can be asked for specific fields in one step. The cost is that they run as API calls or on a GPU, they can hallucinate plausible text where the image is unreadable, and they need validation downstream.

Also asked as: ocr with llm · vision language model ocr · llm for document extraction · gpt for ocr · ocr free document understanding · donut model · layoutlm · best ai ocr · ai ocr vs traditional ocr

LayoutLM pre-trained on text plus its position on the page and set the pattern for document understanding [5]; Donut showed that a model could go straight from pixels to structured output with no OCR engine at all [6]. Mistral's document API is one of several commercial options that return page text as Markdown with tables and images preserved [7], and it is what the OCR & Speech Workspace uses for its page batches [12].

The hallucination row is the one to remember. A classic engine produces garbage on a bad scan and you can see it is garbage. A language model produces a fluent, wrong invoice total. Validation is not optional.

How do I OCR a PDF for free?

If the PDF already has text, extract it with PyMuPDF or a similar library and skip OCR [11]. If it is a scan, run OCRmyPDF, which uses Tesseract to add an invisible text layer so the PDF becomes searchable and copyable [10]. For photos or messy pages, PaddleOCR or EasyOCR in a short Python script. For the best quality on hard documents at some cost, a vision-language document API. All of the open tools run offline on a laptop.

Also asked as: how to ocr a pdf · how to ocr a pdf for free · how to make a scanned pdf searchable · ocr pdf python · free ocr software · best free ocr for pdf · convert scanned pdf to text · how to extract text from scanned pdf · ocr online free

The check-then-OCR pattern in Python:

import fitz  # PyMuPDF

doc = fitz.open("scan.pdf")
for page in doc:
    text = page.get_text()
    if text.strip():
        use(text, page.number + 1)          # already has a text layer
    else:
        pix = page.get_pixmap(dpi=300)       # render, then OCR this page only
        text = ocr_engine(pix.tobytes("png"))
        use(text, page.number + 1)

Keep the page number with every span from the first line. Every citation, search result and audit trail later depends on it.

How do I OCR large PDFs quickly?

Split by page, process pages concurrently in batches sized to your engine's or API's rate limit, and merge in page order. A 500-page manual is 500 independent jobs; a concurrency of eight turns an hour into minutes. Cache results by document hash and page, so a re-upload does not re-OCR. Store text per page, not per document, so retrieval and citations stay page-precise.

Also asked as: ocr large pdf · batch ocr pdf · ocr many pages fast · parallel ocr python · ocr performance · ocr at scale · how long does ocr take

Page-batched concurrency is the design of the OCR & Speech Workspace: pages fan out to the document API in batches, results land in a per-page store with the page number, and the document becomes a scoped knowledge base whose answers cite exact pages [12].

Requirements gathering and user flows, on stage
Requirements gathering and user flows, on stage

How accurate is OCR, and how do I improve it?

On clean 300 DPI scans of printed English, modern engines exceed 99 percent character accuracy. Accuracy falls with low resolution, skew, shadows, coloured backgrounds, unusual fonts, handwriting and dense tables, and it falls hardest on the characters that matter most in business documents: digits. Improve it by fixing the input first, resolution, deskew, contrast, then choosing the right engine for the page type, then validating with rules, and only then by trying a bigger model.

Also asked as: ocr accuracy · how to improve ocr accuracy · why ocr is not accurate · ocr accuracy on handwriting · ocr for low quality images · ocr digits accuracy · ocr confidence score

What is document AI, and how is it different from OCR?

Document AI is OCR followed by understanding: classifying the document type, locating the fields that matter, extracting them into a schema you defined, validating them, and returning confidence per field. OCR gives you a wall of text. Document AI gives you an invoice number, a vendor, line items and a total, typed and checked. The extraction may use a language model with a strict output schema, a layout model trained on forms, or both.

Also asked as: what is document ai · document ai vs ocr · intelligent document processing · idp vs ocr · document understanding ai · document extraction ai · what is idp · ai document processing

DocuLens AI is the field-and-validation layer: extraction through a language model constrained to a JSON schema, then deterministic checks and a compliance score computed the same way every time, so finance and legal users can rely on it [13].

How do I extract invoice data I can trust?

Define the schema first: the fields, their types, and the rules between them. Extract with a model constrained to that schema, or with a layout model trained on invoices. Then validate deterministically: line items must sum to the subtotal, tax must match the rate, dates must parse, the invoice number must match the vendor's pattern, currency must be stated. Flag anything that fails for a human. Score each document the same way every time. Never let the model be the last word on a number.

Also asked as: invoice ocr · how to extract data from invoices · invoice extraction ai · invoice data extraction python · ocr for invoices · extract fields from invoice · invoice parsing llm · receipt ocr

Determinism matters because the same invoice must produce the same result on Tuesday as on Monday. That is why DocuLens computes scores from rules, not from a model's opinion [13].

Can OCR read handwriting?

Partly. Handwriting recognition on clean, separated block letters works reasonably with modern models; cursive, mixed scripts and forms with lines through the text remain hard. Vision-language models do better than classic engines, and confidence scores plus human review are essential. For Indian forms, mixed Hindi and English handwriting on printed templates, the practical approach is to recognise the printed template's fields, then OCR each handwritten field region separately with the expected type as a constraint, and route low-confidence results to a person.

Also asked as: can ocr read handwriting · handwriting ocr · handwriting recognition ai · ocr for handwritten forms · icr vs ocr · handwriting to text ai · ocr hindi handwriting

Which OCR should I use: Tesseract, PaddleOCR, EasyOCR, or an API?

Tesseract via OCRmyPDF for clean scanned documents at zero cost [1][10]. PaddleOCR when you need tables, multilingual pages and production quality on your own hardware [8]. EasyOCR for a quick script [9]. A vision-language document API when documents are messy, layouts matter and you can pay per page and validate the output [7]. Most real pipelines use two: a cheap engine first, an expensive model only on pages the cheap one fails.

Also asked as: best ocr · best ocr api · best ocr for python · best ocr for invoices · best ocr for tables · best ocr 2026 · paddleocr vs tesseract vs easyocr · google vision ocr vs tesseract · azure ocr vs aws textract

What are common OCR and document AI interview questions?

Explain the three stages and where errors come from. Compare classic OCR with vision-language models, including the hallucination risk. Describe how you would OCR a 1,000-page PDF quickly. Explain how you would validate an extracted invoice. Describe how page-level citations work in a document chatbot. If you have built one extraction pipeline with validation, you can answer all of them.

Also asked as: ocr interview questions · document ai interview questions · computer vision interview questions ocr · idp interview questions

Where should I start with an OCR project?

Take twenty real documents you have, invoices, marksheets, notes, and build the check-then-OCR script above with Tesseract. Measure character accuracy on five pages by hand. Then swap in a vision-language model for the pages that fail and add three validation rules. You will have learned every concept on this page, and the pipeline will be the shape of a real product. My longer design note on the pipeline is on the portfolio [14]. For a hands-on document AI session at your college, email pranjulrathour41@gmail.com or use pranjulrathour.scult.in/invite.

Sources

  1. Tesseract OCR, open-source enginegithub.com
  2. Smith, An Overview of the Tesseract OCR Engine (ICDAR 2007)ieeexplore.ieee.org
  3. Baek et al., Character Region Awareness for Text Detection (CRAFT, 2019)arxiv.org
  4. Li et al., TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models (2021)arxiv.org
  5. Xu et al., LayoutLM: Pre-training of Text and Layout for Document Image Understanding (2020)arxiv.org
  6. Kim et al., OCR-free Document Understanding Transformer (Donut, 2022)arxiv.org
  7. Mistral AI, Document AI and OCR documentationdocs.mistral.ai
  8. PaddleOCR, open-source OCR toolkitgithub.com
  9. EasyOCR, ready-to-use OCR in 80+ languagesgithub.com
  10. OCRmyPDF: add an OCR text layer to scanned PDFsocrmypdf.readthedocs.io
  11. PyMuPDF documentationpymupdf.readthedocs.io
  12. OCR & Speech Workspace source code, Pranjul Rathourgithub.com
  13. DocuLens AI source code, Pranjul Rathourgithub.com
  14. Designing an OCR pipeline from scanned PDF to searchable text, Pranjul Rathourpranjulrathour.scult.in
Pranjul Rathour
Pranjul Rathour
GenAI Engineer · Kanpur, Uttar Pradesh, India

GenAI engineer and AI product builder with 2+ years shipping production-grade AI systems: RAG pipelines, fine-tuned LLMs, hybrid retrieval and multi-modal apps across vision, speech and OCR, architected end to end from ingestion to deployment. Leads engineering for SCULT INDIA's 14-member team, founded the 500+ member TechVerse Enclave community and has mentored 200+ students. Three hackathon first prizes: Changethon 2025 (IIT Roorkee), Product Genesis at Vividhotsava 2025 (CSJMU Kanpur) and BYTEBATTLE (MeetKats).

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges. Email pranjulrathour41@gmail.com.

Keep reading

More on Vision, speech & OCR

All Vision, speech & OCR guides →
Campus talks · hackathon judging · mentoring

Want this as a live session at your college?

Open to GenAI roles, hackathon judging, mentorship sessions and guest talks at colleges.

Pranjul Rathour, GenAI engineer in Kanpur