LlamaIndex 2026 document parsing benchmark. ~2,078 human-verified pages from ~1,211 enterprise documents (insurance, finance, government) with 169K rule-based tests across five dimensions: tables (GTRM), charts (ChartDataPointMatch), content faithfulness, semantic formatting, and visual grounding. No LLM-as-judge. Overall score = unweighted mean of the five dimensions.
Accuracy is the reported evaluation metric for ParseBench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | LlamaParse Agentic | verified | 84.9 | 2026 | Source ↗ | Looks wrong? |
| 02 | LlamaParse Cost Effective | verified | 71.9 | 2026 | Source ↗ | Looks wrong? |
| 03 | Google Gemini 3 Flash | verified | 71 | 2026 | Source ↗ | Looks wrong? |
| 04 | Reducto | verified | 67.8 | 2026 | Source ↗ | Looks wrong? |
| 05 | Qwen 3 VL | verified | 62 | 2026 | Source ↗ | Looks wrong? |
| 06 | Azure Document Intelligence | verified | 59.6 | 2026 | Source ↗ | Looks wrong? |
| 07 | Extend | verified | 55.8 | 2026 | Source ↗ | Looks wrong? |
| 08 | Dots OCR 1.5 | verified | 55.8 | 2026 | Source ↗ | Looks wrong? |
| 09 | Docling | verified | 50.6 | 2026 | Source ↗ | Looks wrong? |
| 10 | Google Cloud Document AI | verified | 50.4 | 2026 | Source ↗ | Looks wrong? |
| 11 | AWS Textract | verified | 47.9 | 2026 | Source ↗ | Looks wrong? |
| 12 | OpenAI GPT-5 Mini | verified | 46.8 | 2026 | Source ↗ | Looks wrong? |
| 13 | LandingAI | verified | 45.2 | 2026 | Source ↗ | Looks wrong? |
| 14 | Anthropic Haiku 4.5 | verified | 45.2 | 2026 | Source ↗ | Looks wrong? |