LLM evaluation · reviewed October 7, 2026

Choose a model with evidence

Use current release documentation to build a shortlist, then compare matched evaluations and your own tasks. The registry includes older and undated records; its refresh time does not make those results current.

Current candidates Recorded evidence
Official documentation checked October 7, 2026

Current models to evaluate

Start with these source-checked candidates. This is a selection guide, not a shared benchmark ranking: the providers use different tasks, agent scaffolds and evaluation settings.

Claude Opus 5.5

A candidate for demanding coding and agent work. Anthropic reports Terminal-Bench 4.0 and FrontierCode v1.1 results; those are different evaluations from the older SWE-bench rows.

Anthropic release and evaluation report ↗

GPT-6 Astra

OpenAI’s highest-intelligence family for demanding reasoning, coding and professional work. Evaluate its cost and latency on your own tasks.

OpenAI current model guide ↗

GPT-6.1 Sol

OpenAI’s balanced option for complex coding, computer use and professional work, with lower cost than Astra. Select the reasoning effort explicitly.

OpenAI model documentation ↗

Gemini 3.8 Flash

Google’s generally available option for long-horizon software engineering and autonomous agents. Check the supported tools and thinking settings for your workflow.

Google Gemini model guide ↗

Newer published coding evidence

Anthropic’s September 22, 2026 Opus 5.5 report lists 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1 (Main). These are publisher-reported results, not CodeSOTA reproductions or updates to old SWE-bench scores.

The report uses different reasoning efforts across evaluations and describes safety fallbacks. Read the evaluation details before comparing systems. Original report and conditions ↗

Older scores retain their original model identity. We have not transferred Opus 4.7 or GPT-5 scores to newer releases. Documentation review dates describe this article, not the age of a benchmark result.

Historical registry coverage

Recorded benchmark evidence

Records below use each dataset’s primary metric and retain the stored model names. They are unverified historical registry claims, listed alphabetically with no overall winner. Source links and dates have not all been checked against the original reports. The registry also lacks consistent benchmark versions, agent scaffolds and tool budgets.

MMLU

The original benchmark uses four-choice questions across 57 subjects. It is separate from MMLU-Pro and should retain its original evaluation settings. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Apertus-70B65.2accuracyResult date unknownSource ↗
Apertus-70B-Instruct69.6accuracyResult date unknownSource ↗
Aria73.3accuracyResult date unknownSource ↗
BitNet b1.58 2B4T53.17accuracyResult date unknownSource ↗
BLT-Entropy 8B57.4accuracyResult date unknownSource ↗
Chameleon 34B65.8accuracyResult date unknownSource ↗
Claude 3.5 Sonnet88.3accuracyResult date unknownSource ↗
Claude 3 Opus86.8accuracyResult date unknownSource ↗
Claude Opus 488.8accuracyResult date unknownSource ↗
Claude Opus 4.591.8accuracy2026-01-01Source ↗
Claude Opus 4.591.6accuracy2025-11-24Source ↗
Claude Opus 4.691.2accuracy2026-03-01Source ↗

Showing 12 of 64 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

MMLU-Pro

A harder knowledge and reasoning suite with up to ten answer choices. Prompting and scoring settings still affect reported results. Benchmark documentation ↗

No primary-metric records available in this view.

· Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

GPQA

Graduate-level science questions. Identify the exact subset, such as Diamond, and the sampling and tool settings before comparing records. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Claude 3.5 Sonnet59.4accuracyResult date unknownSource ↗
Claude 3 Opus50.4accuracyResult date unknownSource ↗
Claude Opus 476.7accuracyResult date unknownSource ↗
Claude Opus 4.574.9accuracyResult date unknownSource ↗
Claude Opus 4.691.3accuracyResult date unknownSource ↗
Claude Sonnet 470accuracyResult date unknownSource ↗
Claude Sonnet 4.689.9accuracyResult date unknownSource ↗
DeepSeek R171.5accuracyResult date unknownSource ↗
DeepSeek-V3.282.4accuracyResult date unknownSource ↗
DeepSeek-V3.2-Speciale85.7accuracyResult date unknownSource ↗
DeepSeek-V4-Flash Max88.1accuracyResult date unknownSource ↗
DeepSeek-V4-Pro Max90.1accuracyResult date unknownSource ↗

Showing 12 of 74 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

AIME 2025

Competition mathematics from the 2025 exam. Keep it separate from AIME 2026 and report attempts, answer extraction and any tool access. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Claude Opus 4.580accuracyResult date unknownSource ↗
DeepSeek R172accuracyResult date unknownSource ↗
DeepSeek-V3.293.1accuracyResult date unknownSource ↗
DeepSeek-V3.2-Speciale96accuracyResult date unknownSource ↗
Gemini 2.5 Flash72accuracyResult date unknownSource ↗
Gemini 2.5 Pro88accuracyResult date unknownSource ↗
Gemini 2.5 Pro86.7accuracyResult date unknownSource ↗
Intern-S1-Pro93.1accuracyResult date unknownSource ↗
Kimi-K2.596.1accuracyResult date unknownSource ↗
NVIDIA-Nemotron-3-Nano-30B-A3B-BF1689.1accuracyResult date unknownSource ↗
o386.7accuracyResult date unknownSource ↗
o4-mini92.7accuracyResult date unknownSource ↗

Showing 12 of 22 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

LiveCodeBench

Contest programming with changing problem windows. Release, date window and pass@k must match. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Claude Opus 457.8pass@12024-03-12Source ↗
Claude Sonnet 452.8pass@12024-03-12Source ↗
Codestral 22B29.5pass@12024-03-12Source ↗
DeepSeek-Coder-V2-Instruct43.4pass@12024-03-12Source ↗
DeepSeek R165.9pass@1Result date unknownSource ↗
DeepSeek-R1-052873.3pass@1Result date unknownSource ↗
DeepSeek-R1-Distill-Llama-70B65.2pass@1Result date unknownSource ↗
DeepSeek-R1-Distill-Qwen-32B62.1pass@1Result date unknownSource ↗
DeepSeek-V349.2pass@12024-03-12Source ↗
DeepSeek-v3-032449.2pass@1Result date unknownSource ↗
Gemini 2.5 Flash63.9pass@1Result date unknownSource ↗
Gemini 2.5 Pro75.6pass@1Result date unknownSource ↗

Showing 12 of 30 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

LiveCodeBench Pro

An Elo-based evaluation. Elo and classic LiveCodeBench pass rates are different metrics and are never averaged together here. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Claude Sonnet 4.51,412eloResult date unknownSource ↗
DeepSeek R11,161eloResult date unknownSource ↗
Gemini 2.5 Flash1,288eloResult date unknownSource ↗
Gemini 2.5 Pro1,769eloResult date unknownSource ↗
Gemini 3.1 Pro2,887eloResult date unknownSource ↗
Gemini 3 Pro2,439eloResult date unknownSource ↗
GPT-52,176eloResult date unknownSource ↗
o31,010eloResult date unknownSource ↗
o4-mini2,092eloResult date unknownSource ↗
Qwen3-235B-A22B1,673eloResult date unknownSource ↗

Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

τ²-bench

Interactive agent evaluation. Preserve the environment, user simulator, tools and run protocol. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Claude 3.7 Sonnet47pass_rateResult date unknownSource URL not recorded
Claude Opus 4.579pass_rate2025-11-24Source ↗
Claude Sonnet 4.563pass_rate2025-09-29Source ↗
Gemini 2.5 Pro54pass_rateResult date unknownSource URL not recorded
Gemini 3 Pro69pass_rate2025-11-18Source ↗
GPT-4o36pass_rateResult date unknownSource URL not recorded
GPT-5.159pass_rateResult date unknownSource URL not recorded
GPT-5.273pass_rate2025-12-11Source ↗

Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

Humanity’s Last Exam

A difficult multidisciplinary evaluation. Text-only, multimodal and tool-assisted runs must be distinguished. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Claude 3.5 Sonnet4.08accuracyResult date unknownSource ↗
Claude 3.7 Sonnet8.04accuracyResult date unknownSource ↗
Claude 4.5 Sonnet13.7accuracyResult date unknownSource URL not recorded
Claude Opus 410.72accuracyResult date unknownSource ↗
Claude Opus 4.111.52accuracyResult date unknownSource ↗
Claude Opus 4.525.2accuracyResult date unknownSource ↗
Claude Opus 4.634.44accuracyResult date unknownSource ↗
Claude Opus 4.619accuracyResult date unknownSource ↗
Claude Opus 4.736.2accuracyResult date unknownSource ↗
Claude Sonnet 47.76accuracyResult date unknownSource ↗
Claude Sonnet 4.513.72accuracyResult date unknownSource ↗
Claude Sonnet 4.613.2accuracyResult date unknownSource ↗

Showing 12 of 74 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

A score needs a protocol

These tables do not claim that every record was independently reproduced. Source URLs, result dates and protocol details are sometimes missing. We show those gaps instead of replacing a missing date with today or filtering a fixed list of older models into a “current frontier.”

There is no composite leaderboard here: knowledge percentages, coding pass rates, agent success rates and Elo describe different measurements.

Coding guide → · Benchmark catalogue → · Methodology →