7,010 unit tests across 1,402 PDF documents. Tests parsing of tables, math, multi-column layouts, old scans, and more.
Base is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | chandra-ocr-0.1.0 | vendor | 99.9 | 2026 | Source ↗ | Looks wrong? |
| 02 | olmocr-v0.4.0 | vendor | 99.7 | 2026 | Source ↗ | Looks wrong? |
| 03 | LightOnOCR-2-1B | vendor | 99.6 | 2026 | Source ↗ | Looks wrong? |
| 04 | Qianfan-OCR | vendor | 99.6 | 2026 | Source ↗ | Looks wrong? |
Headers Footers is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | olmocr-v0.4.0 | vendor | 96.1 | 2026 | Source ↗ | Looks wrong? |
| 02 | olmocr-v0.3.0 | vendor | 95.1 | 2026 | Source ↗ | Looks wrong? |
| 03 | chandra-ocr-0.1.0 | vendor | 90.8 | 2026 | Source ↗ | Looks wrong? |
| 04 | Qianfan-OCR | vendor | 42 | 2026 | Source ↗ | Looks wrong? |
Long Tiny Text is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | chandra-ocr-0.1.0 | vendor | 92.3 | 2026 | Source ↗ | Looks wrong? |
| 02 | LightOnOCR-2-1B | vendor | 91.4 | 2026 | Source ↗ | Looks wrong? |
| 03 | olmocr-v0.4.0 | vendor | 81.9 | 2026 | Source ↗ | Looks wrong? |
| 04 | Qianfan-OCR | vendor | 80.4 | 2026 | Source ↗ | Looks wrong? |
Multi Column is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Qianfan-OCR | vendor | 92.2 | 2026 | Source ↗ | Looks wrong? |
| 02 | LightOnOCR-2-1B | vendor | 84.8 | 2026 | Source ↗ | Looks wrong? |
| 03 | olmocr-v0.4.0 | vendor | 83.7 | 2026 | Source ↗ | Looks wrong? |
| 04 | chandra-ocr-0.1.0 | vendor | 81.2 | 2026 | Source ↗ | Looks wrong? |
Arxiv is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | LightOnOCR-2-1B | vendor | 89.6 | 2026 | Source ↗ | Looks wrong? |
| 02 | marker-1.10.0 | vendor | 83.8 | 2026 | Source ↗ | Looks wrong? |
| 03 | olmocr-v0.4.0 | vendor | 83 | 2026 | Source ↗ | Looks wrong? |
| 04 | chandra-ocr-0.1.0 | vendor | 82.2 | 2026 | Source ↗ | Looks wrong? |
| 05 | Qianfan-OCR | vendor | 80.1 | 2026 | Source ↗ | Looks wrong? |
Tables is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | LightOnOCR-2-1B | vendor | 89 | 2026 | Source ↗ | Looks wrong? |
| 02 | dots-ocr-3b | vendor | 88.3 | 2026 | Source ↗ | Looks wrong? |
| 03 | chandra-ocr-0.1.0 | vendor | 88 | 2026 | Source ↗ | Looks wrong? |
| 04 | olmocr-v0.4.0 | vendor | 84.9 | 2026 | Source ↗ | Looks wrong? |
| 05 | Qianfan-OCR | vendor | 81.6 | 2026 | Source ↗ | Looks wrong? |
Old Scans Math is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | LightOnOCR-2-1B | vendor | 85.6 | 2026 | Source ↗ | Looks wrong? |
| 02 | olmocr-v0.4.0 | vendor | 82.3 | 2026 | Source ↗ | Looks wrong? |
| 03 | chandra-ocr-0.1.0 | vendor | 80.3 | 2026 | Source ↗ | Looks wrong? |
| 04 | olmocr-v0.3.0 | vendor | 79.9 | 2026 | Source ↗ | Looks wrong? |
Pass Rate is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
Old Scans is the reported evaluation metric for olmOCR-Bench. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | Qianfan-OCR | vendor | 73.1 | 2026 | Source ↗ | Looks wrong? |
| 02 | chandra-ocr-0.1.0 | vendor | 50.4 | 2026 | Source ↗ | Looks wrong? |
| 03 | olmocr-v0.4.0 | vendor | 47.7 | 2026 | Source ↗ | Looks wrong? |
| 04 | LightOnOCR-2-1B | vendor | 42.2 | 2026 | Source ↗ | Looks wrong? |
| 05 | gpt-4o | vendor | 40.7 | 2026 | Source ↗ | Looks wrong? |