Codesota · Benchmark · AIME 2024Home/Leaderboards/Language & Knowledge/Mathematical Reasoning/AIME 2024
Unknown

AIME 2024.

30 challenging math problems from the 2024 AIME competition. Tests advanced mathematical reasoning.

Paper ↗Leaderboard ↓Lineage
§ 01 · Leaderboard

Results by metric.

Only 3 models on this benchmark
Help build the community leaderboard — submit your model results.
Found a wrong score or missing run?
Use row edits to send a sourced correction into moderation.
Add / edit result ↗Report issue ↗

accuracy

Accuracy is the reported evaluation metric for AIME 2024. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.

Higher is better

Trust tiers for accuracyverifiedpapervendorcommunityunverified

Muted rows were not state of the art when published — an earlier or same-year result already scored better.

RankModelTrustScoreYearLinksFix
01o1-preview
Non-API entry from src
unverified83.32026N/ALooks wrong?
02claude-35-opus
Non-API entry from src
unverified162026N/ALooks wrong?
03gpt-4o
Non-API entry from src
unverified13.42026N/ALooks wrong?
Lineage

AIME 2024 in context.

See full mathematical reasoning benchmarks lineage →
This benchmark (1)
active2024-03
AIME 2024
Successors (1)
active2024-11
FrontierMath
AIME problems are finite and increasingly contaminated as training sets grow. FrontierMath sources unpublished research-frontier problems — contamination by design impossible. The step change from competition math to research math.
§ 04 · Submit a result

Add to the leaderboard.

Submit a Result

Sign in to submit benchmark results for AIME 2024.

Sign in
← Back to Mathematical Reasoning