Claude Opus 5.5
A candidate for demanding coding and agent work. Anthropic reports Terminal-Bench 4.0 and FrontierCode v1.1 results; those are different evaluations from the older SWE-bench rows.
Anthropic release and evaluation report ↗A model, its tools and the harness together produce an agent result. Compare the same task suite, reasoning settings, execution environment and budget. The former April tables do not establish a current October winner.
Current model candidatesStart with these source-checked candidates. This is a selection guide, not a shared benchmark ranking: the providers use different tasks, agent scaffolds and evaluation settings.
A candidate for demanding coding and agent work. Anthropic reports Terminal-Bench 4.0 and FrontierCode v1.1 results; those are different evaluations from the older SWE-bench rows.
Anthropic release and evaluation report ↗OpenAI’s highest-intelligence family for demanding reasoning, coding and professional work. Evaluate its cost and latency on your own tasks.
OpenAI current model guide ↗OpenAI’s balanced option for complex coding, computer use and professional work, with lower cost than Astra. Select the reasoning effort explicitly.
OpenAI model documentation ↗Google’s generally available option for long-horizon software engineering and autonomous agents. Check the supported tools and thinking settings for your workflow.
Google Gemini model guide ↗Anthropic’s September 22, 2026 Opus 5.5 report lists 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1 (Main). These are publisher-reported results, not CodeSOTA reproductions or updates to old SWE-bench scores.
The report uses different reasoning efforts across evaluations and describes safety fallbacks. Read the evaluation details before comparing systems. Original report and conditions ↗
Older scores retain their original model identity. We have not transferred Opus 4.7 or GPT-5 scores to newer releases. Documentation review dates describe this article, not the age of a benchmark result.
SWE-bench evaluates patches against repository tests. Keep Verified, Pro and other splits separate, and preserve the agent scaffold, tools and retry budget.
SWE-bench official leaderboard ↗Terminal-Bench evaluates tasks in terminal environments. Do not compare releases 2.0 and 4.0 as if they were the same task set. Publisher scores also need their harness and effort settings.
Current publisher evaluation report ↗METR’s horizon is the human-expert task duration at which an agent is predicted to succeed at a given reliability. A 50% horizon is not the time an agent can run without failing. The source may lag model releases; check its measurement date.
METR methodology and measurements ↗Agent success depends on the environment and simulator as well as the model. Pin their versions and inspect traces for invalid tool calls, reward hacks and task failures.
τ²-bench official harness ↗Official harness results, local artifacts and pending requests are different evidence states. Local memory tests do not establish an official leaderboard position.
| Resource | Scope | Status | Evidence |
|---|---|---|---|
| Agent Memory Benchmark (AMB) | Provider harness | Track for official scores | vectorize-io/agent-memory-benchmark ↗ |
| Audrey memory artifacts | Local deterministic evidence | Evidence only; no leaderboard claim | HF report + raw artifacts ↗ |
| Audrey AMB provider request | Evaluation route | Pending official harness run | AMB issue #11 ↗ |
This is evidence coverage, not a measured ranking. The previous BinaryAudit, OTelBench, YC-Bench and approximate METR tables were withdrawn until their exact records and protocols can be verified.
Coding benchmark records → · Research artifacts →