Agent evaluation · reviewed October 7, 2026

Evaluate the whole agent

A model, its tools and the harness together produce an agent result. Compare the same task suite, reasoning settings, execution environment and budget. The former April tables do not establish a current October winner.

Current model candidates
Official documentation checked October 7, 2026

Current models to evaluate

Start with these source-checked candidates. This is a selection guide, not a shared benchmark ranking: the providers use different tasks, agent scaffolds and evaluation settings.

Claude Opus 5.5

A candidate for demanding coding and agent work. Anthropic reports Terminal-Bench 4.0 and FrontierCode v1.1 results; those are different evaluations from the older SWE-bench rows.

Anthropic release and evaluation report ↗

GPT-6 Astra

OpenAI’s highest-intelligence family for demanding reasoning, coding and professional work. Evaluate its cost and latency on your own tasks.

OpenAI current model guide ↗

GPT-6.1 Sol

OpenAI’s balanced option for complex coding, computer use and professional work, with lower cost than Astra. Select the reasoning effort explicitly.

OpenAI model documentation ↗

Gemini 3.8 Flash

Google’s generally available option for long-horizon software engineering and autonomous agents. Check the supported tools and thinking settings for your workflow.

Google Gemini model guide ↗

Newer published coding evidence

Anthropic’s September 22, 2026 Opus 5.5 report lists 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1 (Main). These are publisher-reported results, not CodeSOTA reproductions or updates to old SWE-bench scores.

The report uses different reasoning efforts across evaluations and describes safety fallbacks. Read the evaluation details before comparing systems. Original report and conditions ↗

Older scores retain their original model identity. We have not transferred Opus 4.7 or GPT-5 scores to newer releases. Documentation review dates describe this article, not the age of a benchmark result.

Choose an evaluation for the work

Repository repair

SWE-bench evaluates patches against repository tests. Keep Verified, Pro and other splits separate, and preserve the agent scaffold, tools and retry budget.

SWE-bench official leaderboard ↗

Terminal workflows

Terminal-Bench evaluates tasks in terminal environments. Do not compare releases 2.0 and 4.0 as if they were the same task set. Publisher scores also need their harness and effort settings.

Current publisher evaluation report ↗

Time horizons

METR’s horizon is the human-expert task duration at which an agent is predicted to succeed at a given reliability. A 50% horizon is not the time an agent can run without failing. The source may lag model releases; check its measurement date.

METR methodology and measurements ↗

Interactive tools

Agent success depends on the environment and simulator as well as the model. Pin their versions and inspect traces for invalid tool calls, reward hacks and task failures.

τ²-bench official harness ↗

Memory evaluation coverage

Official harness results, local artifacts and pending requests are different evidence states. Local memory tests do not establish an official leaderboard position.

ResourceScopeStatusEvidence
Agent Memory Benchmark (AMB)Provider harnessTrack for official scoresvectorize-io/agent-memory-benchmark ↗
Audrey memory artifactsLocal deterministic evidenceEvidence only; no leaderboard claimHF report + raw artifacts ↗
Audrey AMB provider requestEvaluation routePending official harness runAMB issue #11 ↗

This is evidence coverage, not a measured ranking. The previous BinaryAudit, OTelBench, YC-Bench and approximate METR tables were withdrawn until their exact records and protocols can be verified.

Coding benchmark records → · Research artifacts →