Claude Opus 5.5
A candidate for demanding coding and agent work. Anthropic reports Terminal-Bench 4.0 and FrontierCode v1.1 results; those are different evaluations from the older SWE-bench rows.
Anthropic release and evaluation report ↗Choose by the work you need to finish: generate a function, solve a contest problem, repair a repository or run an agent. Current model documentation and historical benchmark scores answer different parts of that decision.
Current candidates Benchmark recordsUse these guides to narrow your options, check the evidence, and test your own workload.
Function generation and repository repair need different tests.
Understand coding benchmarksCompare candidates for the coding workflow you use.
Code model guideCheck tools, scaffolds, and execution constraints for agents.
Agent benchmark guideFor repository repair, check the SWE-bench protocol and harness.
SWE-bench explainedUse the Claude Code guide as one implementation path; evaluate on your own tasks.
Claude Code implementation guideStart with these source-checked candidates. This is a selection guide, not a shared benchmark ranking: the providers use different tasks, agent scaffolds and evaluation settings.
A candidate for demanding coding and agent work. Anthropic reports Terminal-Bench 4.0 and FrontierCode v1.1 results; those are different evaluations from the older SWE-bench rows.
Anthropic release and evaluation report ↗OpenAI’s highest-intelligence family for demanding reasoning, coding and professional work. Evaluate its cost and latency on your own tasks.
OpenAI current model guide ↗OpenAI’s balanced option for complex coding, computer use and professional work, with lower cost than Astra. Select the reasoning effort explicitly.
OpenAI model documentation ↗Google’s generally available option for long-horizon software engineering and autonomous agents. Check the supported tools and thinking settings for your workflow.
Google Gemini model guide ↗Anthropic’s September 22, 2026 Opus 5.5 report lists 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1 (Main). These are publisher-reported results, not CodeSOTA reproductions or updates to old SWE-bench scores.
The report uses different reasoning efforts across evaluations and describes safety fallbacks. Read the evaluation details before comparing systems. Original report and conditions ↗
Older scores retain their original model identity. We have not transferred Opus 4.7 or GPT-5 scores to newer releases. Documentation review dates describe this article, not the age of a benchmark result.
Records below use each dataset’s primary metric and retain the stored model names. They are unverified historical registry claims, listed alphabetically with no overall winner. Source links and dates have not all been checked against the original reports. The registry also lacks consistent benchmark versions, agent scaffolds and tool budgets.
Repository issue resolution. Compare the same task split, agent scaffold, tools and attempt budget; a model name alone does not define a run. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Claude 3.5 Haiku | 40.6 | resolve-rate | Result date unknown | Source ↗ |
| Claude 3.5 Sonnet | 50.8 | resolve-rate | Result date unknown | Source ↗ |
| Claude 3.7 Sonnet | 63.7 | resolve-rate | Result date unknown | Source ↗ |
| Claude Haiku 4.5 | 73.3 | resolve-rate | Result date unknown | Source ↗ |
| Claude Opus 4 | 72.5 | resolve-rate | Result date unknown | Source ↗ |
| Claude Opus 4.5 | 80.9 | resolve-rate | 2025-11-24 | Source ↗ |
| Claude Opus 4.6 | 80.8 | resolve-rate | 2026-02-17 | Source ↗ |
| Claude Opus 4.7 | 87.6 | resolve-rate | 2026-04-18 | Source ↗ |
| Claude Sonnet 4 | 72.7 | resolve-rate | Result date unknown | Source ↗ |
| Claude Sonnet 4.5 | 77.2 | resolve-rate | Result date unknown | Source ↗ |
| Claude Sonnet 4.6 | 79.6 | resolve-rate | 2026-02-17 | Source ↗ |
| DeepSeek R1 | 49.2 | resolve-rate | Result date unknown | Source ↗ |
Showing 12 of 39 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
164 Python function-completion problems. Original HumanEval and expanded EvalPlus tests are different evaluations. Report the test suite and pass@k setting. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Claude 3.5 Sonnet | 92 | pass@1 | Result date unknown | Source ↗ |
| Claude 3 Opus | 84.9 | pass@1 | Result date unknown | Source ↗ |
| Claude Opus 4 | 92.2 | pass@1 | Result date unknown | Source ↗ |
| Claude Opus 4.6 | 96.3 | pass@1 | 2026-01-01 | Source ↗ |
| Claude Sonnet 4 | 90.6 | pass@1 | Result date unknown | Source ↗ |
| Claude Sonnet 4.6 | 94.1 | pass@1 | 2026-01-01 | Source ↗ |
| Code Llama 34B | 62.4 | pass@1 | Result date unknown | Source ↗ |
| Codestral 22B | 81.1 | pass@1 | 2024-05-29 | Source ↗ |
| Codestral 25.01 | 85.3 | pass@1 | 2025-01-01 | Source ↗ |
| Codex (davinci-002) | 46.9 | pass@1 | 2021-07-01 | Source ↗ |
| DeepSeek-Coder-33B-Instruct | 79.3 | pass@1 | 2023-11-01 | Source ↗ |
| DeepSeek-Coder-V2-Instruct | 90.2 | pass@1 | 2024-06-17 | Source ↗ |
Showing 12 of 42 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
Contest programming problems collected over time. A comparison must use the same release and problem date window; later windows cannot be merged into one score history. Benchmark documentation ↗
| Recorded model | Stored value | Metric | Registry date (unverified) | Stored source |
|---|---|---|---|---|
| Claude Opus 4 | 57.8 | pass@1 | 2024-03-12 | Source ↗ |
| Claude Sonnet 4 | 52.8 | pass@1 | 2024-03-12 | Source ↗ |
| Codestral 22B | 29.5 | pass@1 | 2024-03-12 | Source ↗ |
| DeepSeek-Coder-V2-Instruct | 43.4 | pass@1 | 2024-03-12 | Source ↗ |
| DeepSeek R1 | 65.9 | pass@1 | Result date unknown | Source ↗ |
| DeepSeek-R1-0528 | 73.3 | pass@1 | Result date unknown | Source ↗ |
| DeepSeek-R1-Distill-Llama-70B | 65.2 | pass@1 | Result date unknown | Source ↗ |
| DeepSeek-R1-Distill-Qwen-32B | 62.1 | pass@1 | Result date unknown | Source ↗ |
| DeepSeek-V3 | 49.2 | pass@1 | 2024-03-12 | Source ↗ |
| DeepSeek-v3-0324 | 49.2 | pass@1 | Result date unknown | Source ↗ |
| Gemini 2.5 Flash | 63.9 | pass@1 | Result date unknown | Source ↗ |
| Gemini 2.5 Pro | 75.6 | pass@1 | Result date unknown | Source ↗ |
Showing 12 of 30 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
Code editing in an agent harness. Verify the language set, edit format and retry configuration in the original report. Benchmark documentation ↗
No primary-metric records available in this view.
· Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.
Record the benchmark release, task window, metric, model version, reasoning effort, tool access and retry budget. A percentage from SWE-bench cannot be compared numerically with Elo from LiveCodeBench Pro.
Use representative issues with known tests. Measure successful changes, regressions, review time, latency and total cost. Preserve prompts and execution traces so another run can reproduce the result.
The previous April supplement and inferred progress charts have been withdrawn: they did not establish an October ranking or consistent evaluation protocol. Older releases remain searchable in the registry under their own names.
Agent benchmark guidance → · LLM evaluation guide → · Evidence methodology →