Codesota · Benchmark · TextVQAHome/Leaderboards/TextVQA
Facebook AI Research

TextVQA.

TextVQA evaluates a model's ability to read and reason about text embedded in images. The test set contains 45,336 questions over 28,408 images with prominent scene text, pushing models beyond pure object recognition into OCR-grounded visual reasoning.

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

Accuracy

VQA-style accuracy across answer variants; higher is better.

Higher is better

Trust tiers for Accuracyverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Qwen2.5-VL 72B
Fetched from CodeSOTA API on 2026-04-20
verified85.52026Source ↗Looks wrong?
02Qwen2-VL 72B
Fetched from CodeSOTA API on 2026-04-20
verified84.92026Source ↗Looks wrong?
03InternVL2-76B
Fetched from CodeSOTA API on 2026-04-20
verified84.42026Source ↗Looks wrong?
04Llama 3.2 Vision 90B
Fetched from CodeSOTA API on 2026-04-20
verified83.42026Source ↗Looks wrong?
05Gemini 1.5 Pro
Fetched from CodeSOTA API on 2026-04-20
verified82.22026Source ↗Looks wrong?
06GPT-4V
Fetched from CodeSOTA API on 2026-04-20
verified782026Source ↗Looks wrong?
07GPT-4o
Fetched from CodeSOTA API on 2026-04-20
verified77.42026Source ↗Looks wrong?
08LLaVA-1.5
Fetched from CodeSOTA API on 2026-04-20
verified61.32026Source ↗Looks wrong?
09BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified42.52026Source ↗Looks wrong?
Lineage

TextVQA in context.

See full visual question answering lineage →
This benchmark (1)
active2019-04
TextVQA
None yet — this is the current frontier.
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for TextVQA.

Sign in
← Back to Leaderboards