Invoice OCR: choose for fields, line items and validation
Evaluate invoice extraction on your schema and suppliers. Text recognition, table reconstruction and accounting field accuracy are different tasks.
Choose by workload
| Approach | Useful starting point | What to validate |
|---|---|---|
| Native PDF extraction | Digital invoices with reliable embedded text. | Reading order, missing glyphs, table cells and mixed scanned pages. |
| PaddleOCR or RapidOCR + field logic | Local text and coordinates with your own schema extraction. | Vendor template changes, tax IDs, dates, decimal separators and multilingual text. |
| Docling or MinerU + schema extraction | Invoices with complex tables or layouts where structured page output helps. | Row associations, merged cells, multi-page continuation and source coordinates. |
| Azure Document Intelligence prebuilt-invoice | A managed invoice-specific field and line-item API. | Supported locales, field coverage, service/API version, region and review rate. |
| Amazon Textract AnalyzeExpense | A managed expense-specific extraction path in an AWS workflow. | Summary fields, line-item groups, normalization and supplier-specific failures. |
| A current vision-language model | Flexible schema extraction for varied documents after a controlled trial. | Exact deployed model ID, image resolution, prompt, invented fields and cost per accepted document. |
Candidate shortlist; row order does not represent an accuracy or speed ranking.
Separate recognition from extraction
Getting every visible character right does not prove that quantity, unit price and VAT were attached to the correct line item. Likewise, a table structure score such as TEDS measures a particular table representation; it cannot establish invoice field accuracy or payment readiness.
Define required fields, nullable fields, line-item structure and evidence locations first. For each returned value, keep the source page and span or crop. Treat absent information as absent rather than filling it from a plausible pattern.
Build the smallest useful held-out invoice set
Include multiple suppliers, currencies, tax layouts, credit notes, long invoices, handwritten changes and phone photos. Hold out entire suppliers or templates so tuning does not leak into the test.
Score exact match for invoice IDs and tax IDs; use declared normalization for dates and amounts. Report field precision/recall, line-item matching, missing documents and review rate. Count retries, model calls and human review when estimating cost. Show results by language and scan quality so a good average cannot conceal an unacceptable failure group.
Managed APIs expose different output contracts
Microsoft documents the prebuilt-invoice model for invoice fields and line items. Amazon documents AnalyzeExpense with summary fields and line-item groups. Compare their normalized values and evidence against the same schema instead of comparing unrelated leaderboard percentages.
Azure invoice model and field schema ↗A local parsing baseline
This Docling example exports page content to Markdown. It is a parsing baseline; it does not extract or validate an invoice schema by itself. Keep the parser output alongside the structured fields to make failures reviewable.
Documentation-based example; no runtime measurement is claimed.
# pip install docling
from docling.document_converter import DocumentConverter
converter = DocumentConverter()
document = converter.convert("invoice.pdf").document
print(document.export_to_markdown())Docling documented conversion API ↗Promote a candidate only after the checks pass
Check that line totals reconcile under your declared rounding rules and that net, tax and gross amounts agree. Flag inconsistencies for review; do not silently rewrite the source. Distinguish extraction confidence from a measured probability of correctness.
Publish the exact checkpoint or API model ID, service version, prompts, input hashes, ground truth, outputs and scoring rules with any reported result. New model releases require new runs. A score from GPT-4o or any other older checkpoint cannot be assigned to a newer model.
Primary sources
Documentation and licensing were checked on 7 October 2026. Pin package versions, model checkpoints and configuration in your own environment; upstream defaults can change.