PDF to text: inspect the text layer, then choose a parser
Born-digital, scanned and mixed PDFs need different handling. Choose plain text, structured page output or OCR according to the document and the output contract.
Choose by workload
| PDF type | First path to test | Quality check |
|---|---|---|
| Born-digital text | pypdf extraction without rasterizing the page. | Reading order, glyph mapping, equations and content that exists only as images. |
| Scanned image pages | Document OCR or a parser configured to OCR. | Missing text, recognition errors, page orientation and resolution. |
| Mixed or previously OCRed pages | Inspect both extracted text and rendered pages; decide per page or region. | A hidden OCR text layer may be wrong, incomplete or unrelated to visible content. |
| Multi-column papers, tables and formulas | Docling, MinerU or another structured document parser. | Reading order, cells, captions, formulas and source locations. |
Candidate shortlist; row order does not represent an accuracy or speed ranking.
Extract a candidate text layer
pypdf extracts existing PDF text; it is not an OCR engine. Its documentation explains why PDF text extraction can be ambiguous. The presence of characters does not prove completeness or correctness, and a fixed character-count threshold cannot distinguish every scan from a digital page.
Use this extraction as a first pass. Compare representative pages with their rendered images, especially tables, formulas, unusual fonts and pages containing scanned regions.
Documentation-based example; no runtime measurement is claimed.
# pip install pypdf
from pypdf import PdfReader
reader = PdfReader("document.pdf")
for page_number, page in enumerate(reader.pages, start=1):
text = page.extract_text() or ""
print(f"--- Page {page_number} ---")
print(text)
if not text.strip():
print("No extracted text: inspect this page and consider OCR.")pypdf extraction behavior and limitations ↗Try a structured conversion when layout matters
Docling converts to a document object and can export Markdown. Configure the chosen OCR and parsing pipeline for scans and mixed pages. A successful export is not a guarantee that every table or formula was preserved.
Documentation-based example; no runtime measurement is claimed.
# pip install docling
from docling.document_converter import DocumentConverter
converter = DocumentConverter()
result = converter.convert("document.pdf")
print(result.document.export_to_markdown())Docling documented conversion and options ↗MinerU offers another documented conversion path
The checked MinerU 4.0 interface supports stateless conversion with an explicit tier. Existing older installations should follow the migration documentation. The runtime and selected models affect results and resource use.
Documentation-based example; no runtime measurement is claimed.
# In an isolated environment, for the checked major-version API:
pip install "mineru>=4.0,<5"
mineru-kit parse document.pdf -o document.md --tier standardMinerU current quickstart ↗Measure the output you will actually consume
For transcription, score CER/WER on declared ground truth. For structured pages, also score reading order, headings, table cells and formulas. OCRBench includes visual text understanding tasks, so its aggregate is not a universal plain-transcription error rate.
For search or RAG, inspect whether chunks keep headings, tables and page references together. Retain original pages and source locations so a downstream answer can be checked. Record package locks, model revisions, settings, failed pages and latency with any published evaluation.
Primary sources
Documentation and licensing were checked on 7 October 2026. Pin package versions, model checkpoints and configuration in your own environment; upstream defaults can change.