Home / OCR / PDF to text: inspect the text layer, then choose a parser
Practical guide · Sources checked 7 October 2026

PDF to text: inspect the text layer, then choose a parser

Born-digital, scanned and mixed PDFs need different handling. Choose plain text, structured page output or OCR according to the document and the output contract.

Evidence status. The previous example called a guessed hosted model interface and treated any short text layer as proof of lossless extraction. Those guarantees and that example have been removed. The examples below follow documented APIs and include an explicit quality check.

Choose by workload

PDF typeFirst path to testQuality check
Born-digital textpypdf extraction without rasterizing the page.Reading order, glyph mapping, equations and content that exists only as images.
Scanned image pagesDocument OCR or a parser configured to OCR.Missing text, recognition errors, page orientation and resolution.
Mixed or previously OCRed pagesInspect both extracted text and rendered pages; decide per page or region.A hidden OCR text layer may be wrong, incomplete or unrelated to visible content.
Multi-column papers, tables and formulasDocling, MinerU or another structured document parser.Reading order, cells, captions, formulas and source locations.

Candidate shortlist; row order does not represent an accuracy or speed ranking.

Extract a candidate text layer

pypdf extracts existing PDF text; it is not an OCR engine. Its documentation explains why PDF text extraction can be ambiguous. The presence of characters does not prove completeness or correctness, and a fixed character-count threshold cannot distinguish every scan from a digital page.

Use this extraction as a first pass. Compare representative pages with their rendered images, especially tables, formulas, unusual fonts and pages containing scanned regions.

Documentation-based example; no runtime measurement is claimed.

# pip install pypdf
from pypdf import PdfReader

reader = PdfReader("document.pdf")
for page_number, page in enumerate(reader.pages, start=1):
    text = page.extract_text() or ""
    print(f"--- Page {page_number} ---")
    print(text)
    if not text.strip():
        print("No extracted text: inspect this page and consider OCR.")
pypdf extraction behavior and limitations ↗

Try a structured conversion when layout matters

Docling converts to a document object and can export Markdown. Configure the chosen OCR and parsing pipeline for scans and mixed pages. A successful export is not a guarantee that every table or formula was preserved.

Documentation-based example; no runtime measurement is claimed.

# pip install docling
from docling.document_converter import DocumentConverter

converter = DocumentConverter()
result = converter.convert("document.pdf")
print(result.document.export_to_markdown())
Docling documented conversion and options ↗

MinerU offers another documented conversion path

The checked MinerU 4.0 interface supports stateless conversion with an explicit tier. Existing older installations should follow the migration documentation. The runtime and selected models affect results and resource use.

Documentation-based example; no runtime measurement is claimed.

# In an isolated environment, for the checked major-version API:
pip install "mineru>=4.0,<5"
mineru-kit parse document.pdf -o document.md --tier standard
MinerU current quickstart ↗

Measure the output you will actually consume

For transcription, score CER/WER on declared ground truth. For structured pages, also score reading order, headings, table cells and formulas. OCRBench includes visual text understanding tasks, so its aggregate is not a universal plain-transcription error rate.

For search or RAG, inspect whether chunks keep headings, tables and page references together. Retain original pages and source locations so a downstream answer can be checked. Record package locks, model revisions, settings, failed pages and latency with any published evaluation.

Primary sources

Documentation and licensing were checked on 7 October 2026. Pin package versions, model checkpoints and configuration in your own environment; upstream defaults can change.

OCR benchmark registry · How we verify results