Claude Opus 5.5
A candidate for demanding coding and agent work. Anthropic reports Terminal-Bench 4.0 and FrontierCode v1.1 results; those are different evaluations from the older SWE-bench rows.
Anthropic release and evaluation report ↗Use current release documentation to build a shortlist, then compare matched evaluations and your own tasks. The registry includes older and undated records; its refresh time does not make those results current.
Current candidates Recorded evidenceStart with these source-checked candidates. This is a selection guide, not a shared benchmark ranking: the providers use different tasks, agent scaffolds and evaluation settings.
A candidate for demanding coding and agent work. Anthropic reports Terminal-Bench 4.0 and FrontierCode v1.1 results; those are different evaluations from the older SWE-bench rows.
Anthropic release and evaluation report ↗OpenAI’s highest-intelligence family for demanding reasoning, coding and professional work. Evaluate its cost and latency on your own tasks.
OpenAI current model guide ↗OpenAI’s balanced option for complex coding, computer use and professional work, with lower cost than Astra. Select the reasoning effort explicitly.
OpenAI model documentation ↗Google’s generally available option for long-horizon software engineering and autonomous agents. Check the supported tools and thinking settings for your workflow.
Google Gemini model guide ↗Anthropic’s September 22, 2026 Opus 5.5 report lists 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1 (Main). These are publisher-reported results, not CodeSOTA reproductions or updates to old SWE-bench scores.
The report uses different reasoning efforts across evaluations and describes safety fallbacks. Read the evaluation details before comparing systems. Original report and conditions ↗
Older scores retain their original model identity. We have not transferred Opus 4.7 or GPT-5 scores to newer releases. Documentation review dates describe this article, not the age of a benchmark result.
Records below use each dataset’s primary metric and retain the stored model names. They are unverified historical registry claims, listed alphabetically with no overall winner. Source links and dates have not all been checked against the original reports. The registry also lacks consistent benchmark versions, agent scaffolds and tool budgets.
The original benchmark uses four-choice questions across 57 subjects. It is separate from MMLU-Pro and should retain its original evaluation settings. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Apertus-70B | 65.2 | accuracy | Result date unknown | Source ↗ |
| Apertus-70B-Instruct | 69.6 | accuracy | Result date unknown | Source ↗ |
| Aria | 73.3 | accuracy | Result date unknown | Source ↗ |
| BitNet b1.58 2B4T | 53.17 | accuracy | Result date unknown | Source ↗ |
| BLT-Entropy 8B | 57.4 | accuracy | Result date unknown | Source ↗ |
| Chameleon 34B | 65.8 | accuracy | Result date unknown | Source ↗ |
| Claude 3.5 Sonnet | 88.3 | accuracy | Result date unknown | Source ↗ |
| Claude 3 Opus | 86.8 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4 | 88.8 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4.5 | 91.8 | accuracy | 2026-01-01 | Source ↗ |
| Claude Opus 4.5 | 91.6 | accuracy | 2025-11-24 | Source ↗ |
| Claude Opus 4.6 | 91.2 | accuracy | 2026-03-01 | Source ↗ |
Showing 12 of 64 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
A harder knowledge and reasoning suite with up to ten answer choices. Prompting and scoring settings still affect reported results. Benchmark documentation ↗
No primary-metric records available in this view.
· Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
Graduate-level science questions. Identify the exact subset, such as Diamond, and the sampling and tool settings before comparing records. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Claude 3.5 Sonnet | 59.4 | accuracy | Result date unknown | Source ↗ |
| Claude 3 Opus | 50.4 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4 | 76.7 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4.5 | 74.9 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4.6 | 91.3 | accuracy | Result date unknown | Source ↗ |
| Claude Sonnet 4 | 70 | accuracy | Result date unknown | Source ↗ |
| Claude Sonnet 4.6 | 89.9 | accuracy | Result date unknown | Source ↗ |
| DeepSeek R1 | 71.5 | accuracy | Result date unknown | Source ↗ |
| DeepSeek-V3.2 | 82.4 | accuracy | Result date unknown | Source ↗ |
| DeepSeek-V3.2-Speciale | 85.7 | accuracy | Result date unknown | Source ↗ |
| DeepSeek-V4-Flash Max | 88.1 | accuracy | Result date unknown | Source ↗ |
| DeepSeek-V4-Pro Max | 90.1 | accuracy | Result date unknown | Source ↗ |
Showing 12 of 74 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
Competition mathematics from the 2025 exam. Keep it separate from AIME 2026 and report attempts, answer extraction and any tool access. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Claude Opus 4.5 | 80 | accuracy | Result date unknown | Source ↗ |
| DeepSeek R1 | 72 | accuracy | Result date unknown | Source ↗ |
| DeepSeek-V3.2 | 93.1 | accuracy | Result date unknown | Source ↗ |
| DeepSeek-V3.2-Speciale | 96 | accuracy | Result date unknown | Source ↗ |
| Gemini 2.5 Flash | 72 | accuracy | Result date unknown | Source ↗ |
| Gemini 2.5 Pro | 88 | accuracy | Result date unknown | Source ↗ |
| Gemini 2.5 Pro | 86.7 | accuracy | Result date unknown | Source ↗ |
| Intern-S1-Pro | 93.1 | accuracy | Result date unknown | Source ↗ |
| Kimi-K2.5 | 96.1 | accuracy | Result date unknown | Source ↗ |
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 | 89.1 | accuracy | Result date unknown | Source ↗ |
| o3 | 86.7 | accuracy | Result date unknown | Source ↗ |
| o4-mini | 92.7 | accuracy | Result date unknown | Source ↗ |
Showing 12 of 22 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
Contest programming with changing problem windows. Release, date window and pass@k must match. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Claude Opus 4 | 57.8 | pass@1 | 2024-03-12 | Source ↗ |
| Claude Sonnet 4 | 52.8 | pass@1 | 2024-03-12 | Source ↗ |
| Codestral 22B | 29.5 | pass@1 | 2024-03-12 | Source ↗ |
| DeepSeek-Coder-V2-Instruct | 43.4 | pass@1 | 2024-03-12 | Source ↗ |
| DeepSeek R1 | 65.9 | pass@1 | Result date unknown | Source ↗ |
| DeepSeek-R1-0528 | 73.3 | pass@1 | Result date unknown | Source ↗ |
| DeepSeek-R1-Distill-Llama-70B | 65.2 | pass@1 | Result date unknown | Source ↗ |
| DeepSeek-R1-Distill-Qwen-32B | 62.1 | pass@1 | Result date unknown | Source ↗ |
| DeepSeek-V3 | 49.2 | pass@1 | 2024-03-12 | Source ↗ |
| DeepSeek-v3-0324 | 49.2 | pass@1 | Result date unknown | Source ↗ |
| Gemini 2.5 Flash | 63.9 | pass@1 | Result date unknown | Source ↗ |
| Gemini 2.5 Pro | 75.6 | pass@1 | Result date unknown | Source ↗ |
Showing 12 of 30 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
An Elo-based evaluation. Elo and classic LiveCodeBench pass rates are different metrics and are never averaged together here. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Claude Sonnet 4.5 | 1,412 | elo | Result date unknown | Source ↗ |
| DeepSeek R1 | 1,161 | elo | Result date unknown | Source ↗ |
| Gemini 2.5 Flash | 1,288 | elo | Result date unknown | Source ↗ |
| Gemini 2.5 Pro | 1,769 | elo | Result date unknown | Source ↗ |
| Gemini 3.1 Pro | 2,887 | elo | Result date unknown | Source ↗ |
| Gemini 3 Pro | 2,439 | elo | Result date unknown | Source ↗ |
| GPT-5 | 2,176 | elo | Result date unknown | Source ↗ |
| o3 | 1,010 | elo | Result date unknown | Source ↗ |
| o4-mini | 2,092 | elo | Result date unknown | Source ↗ |
| Qwen3-235B-A22B | 1,673 | elo | Result date unknown | Source ↗ |
Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
Interactive agent evaluation. Preserve the environment, user simulator, tools and run protocol. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Claude 3.7 Sonnet | 47 | pass_rate | Result date unknown | Source URL not recorded |
| Claude Opus 4.5 | 79 | pass_rate | 2025-11-24 | Source ↗ |
| Claude Sonnet 4.5 | 63 | pass_rate | 2025-09-29 | Source ↗ |
| Gemini 2.5 Pro | 54 | pass_rate | Result date unknown | Source URL not recorded |
| Gemini 3 Pro | 69 | pass_rate | 2025-11-18 | Source ↗ |
| GPT-4o | 36 | pass_rate | Result date unknown | Source URL not recorded |
| GPT-5.1 | 59 | pass_rate | Result date unknown | Source URL not recorded |
| GPT-5.2 | 73 | pass_rate | 2025-12-11 | Source ↗ |
Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
A difficult multidisciplinary evaluation. Text-only, multimodal and tool-assisted runs must be distinguished. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Claude 3.5 Sonnet | 4.08 | accuracy | Result date unknown | Source ↗ |
| Claude 3.7 Sonnet | 8.04 | accuracy | Result date unknown | Source ↗ |
| Claude 4.5 Sonnet | 13.7 | accuracy | Result date unknown | Source URL not recorded |
| Claude Opus 4 | 10.72 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4.1 | 11.52 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4.5 | 25.2 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4.6 | 34.44 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4.6 | 19 | accuracy | Result date unknown | Source ↗ |
| Claude Opus 4.7 | 36.2 | accuracy | Result date unknown | Source ↗ |
| Claude Sonnet 4 | 7.76 | accuracy | Result date unknown | Source ↗ |
| Claude Sonnet 4.5 | 13.72 | accuracy | Result date unknown | Source ↗ |
| Claude Sonnet 4.6 | 13.2 | accuracy | Result date unknown | Source ↗ |
Showing 12 of 74 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
These tables do not claim that every record was independently reproduced. Source URLs, result dates and protocol details are sometimes missing. We show those gaps instead of replacing a missing date with today or filtering a fixed list of older models into a “current frontier.”
There is no composite leaderboard here: knowledge percentages, coding pass rates, agent success rates and Elo describe different measurements.
Coding guide → · Benchmark catalogue → · Methodology →