mmmu is a state-of-the-art machine learning benchmark indexed on Codesota. This page tracks published model results, top scores per metric, and the SOTA timeline for mmmu.
ScienceQA proved the multimodal-reasoning benchmark concept worked; MMMU raised scope to college-level knowledge across 30 disciplines with genuine exam difficulty. Top models saturating ScienceQA ~90% while scoring 56% on MMMU at launch confirmed the benchmark transition was necessary.
MMMU-Pro was built after evidence emerged that models were exploiting MMMU's 4-option format and text-based shortcuts. 10-option questions and a vision-only condition expose how much language-mediated pattern matching inflated MMMU scores.
active2024-08
CharXiv
CharXiv narrows focus to scientific chart understanding using real arXiv figures — finer-grained than MMMU's chart questions and sourced from the domain where imprecision in chart reading matters most.