Codesota · Benchmark · MATHHome/Leaderboards/Language & Knowledge/Mathematical Reasoning/MATH
Unknown

MATH.

12,500 competition mathematics problems (5,000 test) from AMC, AIME, and other sources covering algebra, geometry, number theory, and more. Harder than GSM8K. Modern evaluations typically use the MATH-500 representative subset.

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

accuracy

Accuracy is the reported evaluation metric for MATH. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for accuracyverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01o4-mini (high)
Non-API entry from src
unverified98.22026N/ALooks wrong?
02o3 (high)
Non-API entry from src
unverified98.12026N/ALooks wrong?
03o3-mini
Non-API entry from src
unverified97.92026N/ALooks wrong?
04o3
Non-API entry from src
unverified97.82026N/ALooks wrong?
05o4-mini
Non-API entry from src
unverified97.52026N/ALooks wrong?
06DeepSeek-R1
Non-API entry from src
unverified97.32026N/ALooks wrong?
07Gemini 2.5 Pro
Non-API entry from src
unverified97.32026N/ALooks wrong?
08o1
Non-API entry from src
unverified96.42026N/ALooks wrong?
09Claude 3.7 Sonnet
Non-API entry from src
unverified96.22026N/ALooks wrong?
10Kimi k1.5
Non-API entry from src
unverified96.22026N/ALooks wrong?
11DeepSeek-R1-Zero
Non-API entry from src
unverified95.92026N/ALooks wrong?
12DeepSeek-R1-Distill-Llama-70B
Non-API entry from src
unverified94.52026N/ALooks wrong?
13DeepSeek-R1-Distill-Qwen-32B
Non-API entry from src
unverified94.32026N/ALooks wrong?
14DeepSeek-V3-0324
Non-API entry from src
unverified942026N/ALooks wrong?
15QwQ-32B
Non-API entry from src
unverified90.62026N/ALooks wrong?
16deepseek-v3
Non-API entry from src
unverified90.22026N/ALooks wrong?
17o1-mini
Non-API entry from src
unverified902026N/ALooks wrong?
18GPT-4.5 Preview
Non-API entry from src
unverified87.12026N/ALooks wrong?
19o1-preview
Non-API entry from src
unverified85.52026N/ALooks wrong?
20GPT-4.1
Non-API entry from src
unverified82.12026N/ALooks wrong?
21gpt-4o
Non-API entry from src
unverified76.62026N/ALooks wrong?
22Grok 2
Non-API entry from src
unverified76.12026N/ALooks wrong?
23Llama 3.1 405B
Non-API entry from src
unverified73.82026N/ALooks wrong?
24GPT-4 Turbo
Non-API entry from src
unverified73.42026N/ALooks wrong?
25claude-35-sonnet
Non-API entry from src
unverified71.12026N/ALooks wrong?
26gpt-4o-mini
Non-API entry from src
unverified70.22026N/ALooks wrong?
27Llama 3.1 70B
Non-API entry from src
unverified682026N/ALooks wrong?
28gemini-15-pro
Non-API entry from src
unverified67.72026N/ALooks wrong?
29Claude 3 Opus
Non-API entry from src
unverified60.12026N/ALooks wrong?
Lineage

MATH in context.

See full mathematical reasoning benchmarks lineage →
This benchmark (1)
saturating2021-11
MATH
Successors (2)
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for MATH.

Sign in
← Back to Mathematical Reasoning