Codesota · Benchmark · mmmuHome/Leaderboards/mmmu
Unknown

mmmu.

mmmu is a state-of-the-art machine learning benchmark indexed on Codesota. This page tracks published model results, top scores per metric, and the SOTA timeline for mmmu.

Paper ↗Lineage
§ 01 · Leaderboard

Results by metric.

No results yet on this benchmark
Help build the community leaderboard — submit your model results.
Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

No benchmark results available yet for mmmu.

Check back soon as we continue collecting data.

Lineage

MMMU in context.

See full multimodal reasoning benchmarks lineage →
Predecessors (1)
saturating2022-09
ScienceQA
ScienceQA proved the multimodal-reasoning benchmark concept worked; MMMU raised scope to college-level knowledge across 30 disciplines with genuine exam difficulty. Top models saturating ScienceQA ~90% while scoring 56% on MMMU at launch confirmed the benchmark transition was necessary.
This benchmark (1)
active2023-11
MMMU
Successors (3)
active2023-10
MathVista
MathVista fills the specific mathematical-reasoning gap in MMMU — visual geometry, charts, and plots require math skills MMMU's multi-discipline framing doesn't isolate. Complementary rather than competitive.
active2024-09
MMMU-Pro
MMMU-Pro was built after evidence emerged that models were exploiting MMMU's 4-option format and text-based shortcuts. 10-option questions and a vision-only condition expose how much language-mediated pattern matching inflated MMMU scores.
active2024-08
CharXiv
CharXiv narrows focus to scientific chart understanding using real arXiv figures — finer-grained than MMMU's chart questions and sourced from the domain where imprecision in chart reading matters most.
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for mmmu.

Sign in
← Back to Leaderboards