Codesota · Benchmark · ParseBenchHome/Leaderboards/Vision & Documents/Document Parsing/ParseBench
Unknown

ParseBench.

LlamaIndex 2026 document parsing benchmark. ~2,078 human-verified pages from ~1,211 enterprise documents (insurance, finance, government) with 169K rule-based tests across five dimensions: tables (GTRM), charts (ChartDataPointMatch), content faithfulness, semantic formatting, and visual grounding. No LLM-as-judge. Overall score = unweighted mean of the five dimensions.

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

accuracy

Accuracy is the reported evaluation metric for ParseBench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for accuracyverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01LlamaParse Agentic
Fetched from CodeSOTA API on 2026-04-20
verified84.92026Source ↗Looks wrong?
02LlamaParse Cost Effective
Fetched from CodeSOTA API on 2026-04-20
verified71.92026Source ↗Looks wrong?
03Google Gemini 3 Flash
Fetched from CodeSOTA API on 2026-04-20
verified712026Source ↗Looks wrong?
04Reducto
Fetched from CodeSOTA API on 2026-04-20
verified67.82026Source ↗Looks wrong?
05Qwen 3 VL
Fetched from CodeSOTA API on 2026-04-20
verified622026Source ↗Looks wrong?
06Azure Document Intelligence
Fetched from CodeSOTA API on 2026-04-20
verified59.62026Source ↗Looks wrong?
07Extend
Fetched from CodeSOTA API on 2026-04-20
verified55.82026Source ↗Looks wrong?
08Dots OCR 1.5
Fetched from CodeSOTA API on 2026-04-20
verified55.82026Source ↗Looks wrong?
09Docling
Fetched from CodeSOTA API on 2026-04-20
verified50.62026Source ↗Looks wrong?
10Google Cloud Document AI
Fetched from CodeSOTA API on 2026-04-20
verified50.42026Source ↗Looks wrong?
11AWS Textract
Fetched from CodeSOTA API on 2026-04-20
verified47.92026Source ↗Looks wrong?
12OpenAI GPT-5 Mini
Fetched from CodeSOTA API on 2026-04-20
verified46.82026Source ↗Looks wrong?
13LandingAI
Fetched from CodeSOTA API on 2026-04-20
verified45.22026Source ↗Looks wrong?
14Anthropic Haiku 4.5
Fetched from CodeSOTA API on 2026-04-20
verified45.22026Source ↗Looks wrong?
Lineage

ParseBench in context.

See full ocr benchmarks lineage →
This benchmark (1)
active2026-01
ParseBench
None yet — this is the current frontier.
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for ParseBench.

Sign in
← Back to Document Parsing