Codesota · Benchmark · VQA v2.0Home/Leaderboards/Multimodal Media/Visual Question Answering/VQA v2.0
Unknown

VQA v2.0.

265K images with 1.1M questions. Balanced dataset to reduce language biases found in v1.

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

accuracy

Accuracy is the reported evaluation metric for VQA v2.0. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for accuracyverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01Qwen2-VL 72B
Fetched from CodeSOTA API on 2026-04-20
verified87.62026Source ↗Looks wrong?
02InternVL2-76B
Fetched from CodeSOTA API on 2026-04-20
verified87.22026Source ↗Looks wrong?
03Gemini 1.5 Pro
Fetched from CodeSOTA API on 2026-04-20
verified86.52026Source ↗Looks wrong?
04PaLI-X 55B
Fetched from CodeSOTA API on 2026-04-20
verified86.12026Source ↗Looks wrong?
05NVLM-D 1.0 72B
Fetched from CodeSOTA API on 2026-04-20
verified85.42026Source ↗Looks wrong?
06NVLM-X 1.0 72B
Fetched from CodeSOTA API on 2026-04-20
verified85.22026Source ↗Looks wrong?
07NVLM-H 1.0 72B
Fetched from CodeSOTA API on 2026-04-20
verified85.22026Source ↗Looks wrong?
08VILA-1.5 40B
Fetched from CodeSOTA API on 2026-04-20
verified84.32026Source ↗Looks wrong?
09LLaVA-NeXT 34B
Fetched from CodeSOTA API on 2026-04-20
verified83.72026Source ↗Looks wrong?
10LLaVA-NeXT 13B
Fetched from CodeSOTA API on 2026-04-20
verified82.82026Source ↗Looks wrong?
11CogVLM-17B
Fetched from CodeSOTA API on 2026-04-20
verified82.32026Source ↗Looks wrong?
12LLaVA-NeXT 7B (Mistral)
Fetched from CodeSOTA API on 2026-04-20
verified82.22026Source ↗Looks wrong?
13BLIP-2
Fetched from CodeSOTA API on 2026-04-20
verified82.192026Source ↗Looks wrong?
14LLaVA-NeXT 7B (Vicuna)
Fetched from CodeSOTA API on 2026-04-20
verified81.82026Source ↗Looks wrong?
15Pixtral Large
Fetched from CodeSOTA API on 2026-04-20
vendor80.92026Source ↗Looks wrong?
16Llama 3-V 405B
Fetched from CodeSOTA API on 2026-04-20
verified80.22026Source ↗Looks wrong?
17LLaVA-1.5 13B
Fetched from CodeSOTA API on 2026-04-20
verified802026Source ↗Looks wrong?
18LLaVA-1.5
Fetched from CodeSOTA API on 2026-04-20
verified802026Source ↗Looks wrong?
19Llama 3-V 70B
Fetched from CodeSOTA API on 2026-04-20
verified79.12026Source ↗Looks wrong?
20Pixtral-12B
Fetched from CodeSOTA API on 2026-04-20
vendor78.62026Source ↗Looks wrong?
21GPT-4o
Fetched from CodeSOTA API on 2026-04-20
verified78.52026Source ↗Looks wrong?
22Llama 3.2 90B Vision Instruct
Fetched from CodeSOTA API on 2026-04-20
vendor78.12026Source ↗Looks wrong?
23GPT-4V
Fetched from CodeSOTA API on 2026-04-20
verified77.22026Source ↗Looks wrong?
Lineage

VQAv2 in context.

See full visual question answering lineage →
Predecessors (1)
superseded2015-05
VQA
Models answered correctly without looking at the image — VQAv2's balanced pairs force visual grounding.
This benchmark (1)
saturated2017-04
VQAv2
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for VQA v2.0.

Sign in
← Back to Visual Question Answering