Codesota · Benchmark · olmOCR-BenchHome/Leaderboards/Vision & Documents/Document Parsing/olmOCR-Bench
Allen Institute for AI

olmOCR-Bench.

7,010 unit tests across 1,402 PDF documents. Tests parsing of tables, math, multi-column layouts, old scans, and more.

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

Base

Base is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Baseverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01chandra-ocr-0.1.0
Fetched from CodeSOTA API on 2026-04-20
vendor99.92026Source ↗Looks wrong?
02olmocr-v0.4.0
Fetched from CodeSOTA API on 2026-04-20
vendor99.72026Source ↗Looks wrong?
03LightOnOCR-2-1B
Fetched from CodeSOTA API on 2026-04-20
vendor99.62026Source ↗Looks wrong?
04Qianfan-OCR
Fetched from CodeSOTA API on 2026-04-20
vendor99.62026Source ↗Looks wrong?

Headers Footers

Headers Footers is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Headers Footersverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01olmocr-v0.4.0
Fetched from CodeSOTA API on 2026-04-20
vendor96.12026Source ↗Looks wrong?
02olmocr-v0.3.0
Fetched from CodeSOTA API on 2026-04-20
vendor95.12026Source ↗Looks wrong?
03chandra-ocr-0.1.0
Fetched from CodeSOTA API on 2026-04-20
vendor90.82026Source ↗Looks wrong?
04Qianfan-OCR
Fetched from CodeSOTA API on 2026-04-20
vendor422026Source ↗Looks wrong?

Long Tiny Text

Long Tiny Text is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Long Tiny Textverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01chandra-ocr-0.1.0
Fetched from CodeSOTA API on 2026-04-20
vendor92.32026Source ↗Looks wrong?
02LightOnOCR-2-1B
Fetched from CodeSOTA API on 2026-04-20
vendor91.42026Source ↗Looks wrong?
03olmocr-v0.4.0
Fetched from CodeSOTA API on 2026-04-20
vendor81.92026Source ↗Looks wrong?
04Qianfan-OCR
Fetched from CodeSOTA API on 2026-04-20
vendor80.42026Source ↗Looks wrong?

Multi Column

Multi Column is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Multi Columnverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Qianfan-OCR
Fetched from CodeSOTA API on 2026-04-20
vendor92.22026Source ↗Looks wrong?
02LightOnOCR-2-1B
Fetched from CodeSOTA API on 2026-04-20
vendor84.82026Source ↗Looks wrong?
03olmocr-v0.4.0
Fetched from CodeSOTA API on 2026-04-20
vendor83.72026Source ↗Looks wrong?
04chandra-ocr-0.1.0
Fetched from CodeSOTA API on 2026-04-20
vendor81.22026Source ↗Looks wrong?

Arxiv

Arxiv is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Arxivverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01LightOnOCR-2-1B
Fetched from CodeSOTA API on 2026-04-20
vendor89.62026Source ↗Looks wrong?
02marker-1.10.0
Fetched from CodeSOTA API on 2026-04-20
vendor83.82026Source ↗Looks wrong?
03olmocr-v0.4.0
Fetched from CodeSOTA API on 2026-04-20
vendor832026Source ↗Looks wrong?
04chandra-ocr-0.1.0
Fetched from CodeSOTA API on 2026-04-20
vendor82.22026Source ↗Looks wrong?
05Qianfan-OCR
Fetched from CodeSOTA API on 2026-04-20
vendor80.12026Source ↗Looks wrong?

Tables

Tables is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Tablesverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01LightOnOCR-2-1B
Fetched from CodeSOTA API on 2026-04-20
vendor892026Source ↗Looks wrong?
02dots-ocr-3b
Fetched from CodeSOTA API on 2026-04-20
vendor88.32026Source ↗Looks wrong?
03chandra-ocr-0.1.0
Fetched from CodeSOTA API on 2026-04-20
vendor882026Source ↗Looks wrong?
04olmocr-v0.4.0
Fetched from CodeSOTA API on 2026-04-20
vendor84.92026Source ↗Looks wrong?
05Qianfan-OCR
Fetched from CodeSOTA API on 2026-04-20
vendor81.62026Source ↗Looks wrong?

Old Scans Math

Old Scans Math is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Old Scans Mathverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01LightOnOCR-2-1B
Fetched from CodeSOTA API on 2026-04-20
vendor85.62026Source ↗Looks wrong?
02olmocr-v0.4.0
Fetched from CodeSOTA API on 2026-04-20
vendor82.32026Source ↗Looks wrong?
03chandra-ocr-0.1.0
Fetched from CodeSOTA API on 2026-04-20
vendor80.32026Source ↗Looks wrong?
04olmocr-v0.3.0
Fetched from CodeSOTA API on 2026-04-20
vendor79.92026Source ↗Looks wrong?

Pass Rate

Pass Rate is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Pass Rateverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01infinity-parser2-pro
Mapped from PWC olmOCR-Bench Accuracy.; Overall score reported on the official Hugging Face olmOCR-bench leaderboard for infly/Infinity-Parser2-Pro. Source: Infinity Parser technical report / HF leaderboard entry; metric: Accuracy.; PWC evaluation id 4945; paper: Infinity-Parser2-Pro
verified87.62026Source ↗Looks wrong?
02chandra-2
Mapped from PWC olmOCR-Bench Accuracy.; Chandra OCR 2 (5B params, Qwen3.5 backbone) reported on olmOCR-Bench in the HF model card; Overall score 85.9 +/- 0.8. Sub-categories on the same benchmark: ArXiv 90.2, Old Scans Math 89.3, Tables 89.9, Old Scans 49.8, Headers and Footers 92.5, Multi column 83.5, Long tiny text 92.1, Base 99.6. Source: own benchmarks, reported in the chandra-ocr-2 HF model card.; PWC evaluation id 1165; paper: Chandra OCR 2
verified85.92026Source ↗Looks wrong?
03dots.mocr
Mapped from PWC olmOCR-Bench Accuracy.; Overall score on olmOCR-Bench from the dots.mocr Hugging Face model card (section 1.2). 3B image-text-to-text VLM; page-header and page-footer cells deleted from the result markdown per the card's note; metric source: olmocr plus internal evaluations.; PWC evaluation id 1148; paper: Multimodal OCR: Parse Anything from Documents
verified83.92026Source ↗Looks wrong?
04LightOnOCR-2-1B
Mapped from PWC olmOCR-Bench Accuracy.; Overall score on olmOCR-Bench reported on the Hugging Face model card. Score excludes the Headers/Footers (H&F) sub-test because that category rewards omission rather than transcription, while LightOnOCR-2 is trained for full-page transcription and intentionally preserves headers/footers.; PWC evaluation id 1147; paper: LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
verified83.22026Source ↗Looks wrong?
05chandra-ocr-0.1.0
Fetched from CodeSOTA API on 2026-04-20
vendor83.12026Source ↗Looks wrong?
06chandra
Mapped from PWC olmOCR-Bench Accuracy.; Reported in the datalab-to/chandra Hugging Face model card on olmOCR-Bench. Overall score 83.1 +/- 0.9; sub-category scores from the same table: ArXiv 82.2, Old Scans Math 80.3, Tables 88.0, Old Scans 50.4, Headers and Footers 90.8, Multi column 81.2, Long tiny text 92.3, Base 99.9. Source: own benchmarks.; PWC evaluation id 1327; paper: Chandra
verified83.12026Source ↗Looks wrong?
07infinity-parser-7b
Mapped from PWC olmOCR-Bench Accuracy.; Overall score reported on the official Hugging Face olmOCR-bench leaderboard for infly/Infinity-Parser-7B. Source: Infinity Parser technical report / HF leaderboard entry; metric: Accuracy.; PWC evaluation id 4946; paper: Infinity Parser: Layout Aware Reinforcement Learning for Scanned Document Parsing
verified82.52026Source ↗Looks wrong?
08olmocr-v0.4.0
Fetched from CodeSOTA API on 2026-04-20
vendor82.42026Source ↗Looks wrong?
09olmocr-2-7b-1025-7b
Mapped from PWC olmOCR-Bench Accuracy.; PWC evaluation id 679; paper: olmOCR 2: Unit Test Rewards for Document OCR
verified82.42026Source ↗Looks wrong?
10falcon-ocr
Mapped from PWC olmOCR-Bench Accuracy.; olmOCR-Bench accuracy reported by tiiuae/Falcon-OCR model card (arXiv:2603.27365, Apr 2026). 'Average' across 8 category splits: ArXiv Math 80.5, Base 99.5, Headers/Footers 94.0, Long Tiny Text 78.5, Multi Column 87.1, Old Scans 43.5, Old Scans Math 69.2, Tables 90.3. Inference via Layout + OCR two-stage pipeline (PP-DocLayoutV3 layout detection + Falcon-OCR category-prompted VLM).; PWC evaluation id 1156; paper: Falcon Perception
verified80.32026Source ↗Looks wrong?
11paddleocr-vl
Mapped from PWC olmOCR-Bench Accuracy.; PWC evaluation id 86; paper: PaddleOCR-VL: Boosting Multilingual Document Parsing via a 0.9B Ultra-Compact Vision-Language Model
verified802026Source ↗Looks wrong?
12Qianfan-OCR
Mapped from PWC olmOCR-Bench Accuracy.; End-to-end OCR on olmOCR-Bench; overall 79.8 (Arxiv Math 80.1, Old Scans Math 73.1, Table Tests 81.6, Old Scans 42.0, Multi Column 80.4, Long Tiny Text 89.1, Headers Footers 92.2).; PWC evaluation id 1196; paper: Qianfan-OCR: A Unified End-to-End Model for Document Intelligence
verified79.82026Source ↗Looks wrong?
13Qwen3-VL-4B
Fetched from CodeSOTA API on 2026-04-20
vendor79.22026Source ↗Looks wrong?
14PaddleOCR-VL-1.5
Fetched from CodeSOTA API on 2026-04-20
vendor79.12026Source ↗Looks wrong?
15dots-ocr-3b
Fetched from CodeSOTA API on 2026-04-20
vendor79.12026Source ↗Looks wrong?
16dots-ocr
Mapped from PWC olmOCR-Bench Accuracy.; Overall score reported on the official Hugging Face olmOCR-bench leaderboard for rednote-hilab/dots.ocr. Source: dots.ocr technical report / HF leaderboard entry; metric: Accuracy.; PWC evaluation id 4947; paper: dots.ocr: Multilingual Document Layout Parsing in a Single Vision-Language Model
verified79.12026Source ↗Looks wrong?
17mistral-ocr-3
Fetched from CodeSOTA API on 2026-04-20
vendor782026Source ↗Looks wrong?
18mineru-2.5
Mapped from PWC olmOCR-Bench Accuracy.; PWC evaluation id 93; paper: MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing
verified77.52026Source ↗Looks wrong?
19marker-1.10.0
Fetched from CodeSOTA API on 2026-04-20
vendor76.52026Source ↗Looks wrong?
20deepseek-ocr-2
Mapped from PWC olmOCR-Bench Accuracy.; Overall score reported on the official Hugging Face olmOCR-bench leaderboard for deepseek-ai/DeepSeek-OCR-2. Source: HF leaderboard entry and DeepSeek-OCR-2 arXiv-linked model card; metric: Accuracy.; PWC evaluation id 4948; paper: DeepSeek-OCR 2: Visual Causal Flow
verified76.32026Source ↗Looks wrong?
21marker-1.10.1
Fetched from CodeSOTA API on 2026-04-20
vendor76.12026Source ↗Looks wrong?
22lightonocr-1b-1025
Mapped from PWC olmOCR-Bench Accuracy.; Overall score reported on the official Hugging Face olmOCR-bench leaderboard for lightonai/LightOnOCR-1B-1025. Source note: Headers & Footers category excluded; metric: Accuracy.; PWC evaluation id 4949; paper: LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
verified76.12026Source ↗Looks wrong?
23MonkeyOCR-pro-3B
Fetched from CodeSOTA API on 2026-04-20
vendor75.82026Source ↗Looks wrong?
24deepseek-ocr
Fetched from CodeSOTA API on 2026-04-20
vendor75.72026Source ↗Looks wrong?
25DeepSeek-OCR
Mapped from PWC olmOCR-Bench Accuracy.; Overall score reported on the official Hugging Face olmOCR-bench leaderboard for deepseek-ai/DeepSeek-OCR. Source: olmOCR-Bench GitHub leaderboard entry and DeepSeek-OCR arXiv-linked model card; metric: Accuracy.; PWC evaluation id 4950; paper: DeepSeek-OCR: Contexts Optical Compression
verified75.72026Source ↗Looks wrong?
26olmocr
Mapped from PWC olmOCR-Bench Accuracy.; Paper Table 4, olmOCR-Bench overall unit-test pass rate for Ours (v0.1.75 Anchored). Category scores reported in the paper: AR 74.9, OSM 71.2, TA 71.0, OS 42.2, HF 94.5, MC 78.3, LTT 73.3, Base 98.3.; PWC evaluation id 4979; paper: olmOCR: Unlocking Trillions of Tokens in PDFs with Vision Language Models
verified75.52026Source ↗Looks wrong?
27mineru-2.5
Fetched from CodeSOTA API on 2026-04-20
vendor75.22026Source ↗Looks wrong?
28GLM-OCR
Mapped from PWC olmOCR-Bench Accuracy.; Overall score reported on the official Hugging Face olmOCR-bench leaderboard for zai-org/GLM-OCR. Source note: Headers & Footers category excluded; evaluated via ZAI API; metric: Accuracy.; PWC evaluation id 4951; paper: GLM-OCR Technical Report
verified75.22026Source ↗Looks wrong?
29mistral-ocr-api
Fetched from CodeSOTA API on 2026-04-20
vendor722026Source ↗Looks wrong?
30firered-ocr
Mapped from PWC olmOCR-Bench Accuracy.; Overall score reported on the official Hugging Face olmOCR-bench leaderboard for FireRedTeam/FireRed-OCR. Source note: Headers & Footers category excluded; metric: Accuracy.; PWC evaluation id 4952; paper: FireRed-OCR Technical Report
verified70.22026Source ↗Looks wrong?
31gpt-4o-anchored
Fetched from CodeSOTA API on 2026-04-20
vendor69.92026Source ↗Looks wrong?
32nanonets-ocr2-3b
Fetched from CodeSOTA API on 2026-04-20
vendor69.52026Source ↗Looks wrong?
33gemini-flash-2
Fetched from CodeSOTA API on 2026-04-20
vendor63.82026Source ↗Looks wrong?

Old Scans

Old Scans is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for Old Scansverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Qianfan-OCR
Fetched from CodeSOTA API on 2026-04-20
vendor73.12026Source ↗Looks wrong?
02chandra-ocr-0.1.0
Fetched from CodeSOTA API on 2026-04-20
vendor50.42026Source ↗Looks wrong?
03olmocr-v0.4.0
Fetched from CodeSOTA API on 2026-04-20
vendor47.72026Source ↗Looks wrong?
04LightOnOCR-2-1B
Fetched from CodeSOTA API on 2026-04-20
vendor42.22026Source ↗Looks wrong?
05gpt-4o
Fetched from CodeSOTA API on 2026-04-20
vendor40.72026Source ↗Looks wrong?
Lineage

olmOCR-Bench in context.

See full ocr benchmarks lineage →
This benchmark (1)
active2025-03
olmOCR-Bench
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for olmOCR-Bench.

Sign in
← Back to Document Parsing