Coding · editorial review October 7, 2026

Choose a tool for coding

Choose by the work you need to finish: generate a function, solve a contest problem, repair a repository or run an agent. Current model documentation and historical benchmark scores answer different parts of that decision.

Current candidates Benchmark records

From a task to a tool you can try

Use these guides to narrow your options, check the evidence, and test your own workload.

  1. Define the task

    Function generation and repository repair need different tests.

    Understand coding benchmarks
  2. Build a shortlist

    Compare candidates for the coding workflow you use.

    Code model guide
  3. Compare tradeoffs

    Check tools, scaffolds, and execution constraints for agents.

    Agent benchmark guide
  4. Inspect the evidence

    For repository repair, check the SWE-bench protocol and harness.

    SWE-bench explained
  5. Try your repository

    Use the Claude Code guide as one implementation path; evaluate on your own tasks.

    Claude Code implementation guide
Official documentation checked October 7, 2026

Current models to evaluate

Start with these source-checked candidates. This is a selection guide, not a shared benchmark ranking: the providers use different tasks, agent scaffolds and evaluation settings.

Claude Opus 5.5

A candidate for demanding coding and agent work. Anthropic reports Terminal-Bench 4.0 and FrontierCode v1.1 results; those are different evaluations from the older SWE-bench rows.

Anthropic release and evaluation report ↗

GPT-6 Astra

OpenAI’s highest-intelligence family for demanding reasoning, coding and professional work. Evaluate its cost and latency on your own tasks.

OpenAI current model guide ↗

GPT-6.1 Sol

OpenAI’s balanced option for complex coding, computer use and professional work, with lower cost than Astra. Select the reasoning effort explicitly.

OpenAI model documentation ↗

Gemini 3.8 Flash

Google’s generally available option for long-horizon software engineering and autonomous agents. Check the supported tools and thinking settings for your workflow.

Google Gemini model guide ↗

Newer published coding evidence

Anthropic’s September 22, 2026 Opus 5.5 report lists 66.4% on Terminal-Bench 4.0 and 54.4% on FrontierCode v1.1 (Main). These are publisher-reported results, not CodeSOTA reproductions or updates to old SWE-bench scores.

The report uses different reasoning efforts across evaluations and describes safety fallbacks. Read the evaluation details before comparing systems. Original report and conditions ↗

Older scores retain their original model identity. We have not transferred Opus 4.7 or GPT-5 scores to newer releases. Documentation review dates describe this article, not the age of a benchmark result.

Historical registry coverage

Recorded benchmark evidence

Records below use each dataset’s primary metric and retain the stored model names. They are unverified historical registry claims, listed alphabetically with no overall winner. Source links and dates have not all been checked against the original reports. The registry also lacks consistent benchmark versions, agent scaffolds and tool budgets.

SWE-bench Verified

Repository issue resolution. Compare the same task split, agent scaffold, tools and attempt budget; a model name alone does not define a run. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Claude 3.5 Haiku40.6resolve-rateResult date unknownSource ↗
Claude 3.5 Sonnet50.8resolve-rateResult date unknownSource ↗
Claude 3.7 Sonnet63.7resolve-rateResult date unknownSource ↗
Claude Haiku 4.573.3resolve-rateResult date unknownSource ↗
Claude Opus 472.5resolve-rateResult date unknownSource ↗
Claude Opus 4.580.9resolve-rate2025-11-24Source ↗
Claude Opus 4.680.8resolve-rate2026-02-17Source ↗
Claude Opus 4.787.6resolve-rate2026-04-18Source ↗
Claude Sonnet 472.7resolve-rateResult date unknownSource ↗
Claude Sonnet 4.577.2resolve-rateResult date unknownSource ↗
Claude Sonnet 4.679.6resolve-rate2026-02-17Source ↗
DeepSeek R149.2resolve-rateResult date unknownSource ↗

Showing 12 of 39 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

HumanEval

164 Python function-completion problems. Original HumanEval and expanded EvalPlus tests are different evaluations. Report the test suite and pass@k setting. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Claude 3.5 Sonnet92pass@1Result date unknownSource ↗
Claude 3 Opus84.9pass@1Result date unknownSource ↗
Claude Opus 492.2pass@1Result date unknownSource ↗
Claude Opus 4.696.3pass@12026-01-01Source ↗
Claude Sonnet 490.6pass@1Result date unknownSource ↗
Claude Sonnet 4.694.1pass@12026-01-01Source ↗
Code Llama 34B62.4pass@1Result date unknownSource ↗
Codestral 22B81.1pass@12024-05-29Source ↗
Codestral 25.0185.3pass@12025-01-01Source ↗
Codex (davinci-002)46.9pass@12021-07-01Source ↗
DeepSeek-Coder-33B-Instruct79.3pass@12023-11-01Source ↗
DeepSeek-Coder-V2-Instruct90.2pass@12024-06-17Source ↗

Showing 12 of 42 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

LiveCodeBench

Contest programming problems collected over time. A comparison must use the same release and problem date window; later windows cannot be merged into one score history. Benchmark documentation ↗

Recorded modelStored valueMetricRegistry date (unverified)Stored source
Claude Opus 457.8pass@12024-03-12Source ↗
Claude Sonnet 452.8pass@12024-03-12Source ↗
Codestral 22B29.5pass@12024-03-12Source ↗
DeepSeek-Coder-V2-Instruct43.4pass@12024-03-12Source ↗
DeepSeek R165.9pass@1Result date unknownSource ↗
DeepSeek-R1-052873.3pass@1Result date unknownSource ↗
DeepSeek-R1-Distill-Llama-70B65.2pass@1Result date unknownSource ↗
DeepSeek-R1-Distill-Qwen-32B62.1pass@1Result date unknownSource ↗
DeepSeek-V349.2pass@12024-03-12Source ↗
DeepSeek-v3-032449.2pass@1Result date unknownSource ↗
Gemini 2.5 Flash63.9pass@1Result date unknownSource ↗
Gemini 2.5 Pro75.6pass@1Result date unknownSource ↗

Showing 12 of 30 records. Open the benchmark record → · Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

Aider Polyglot

Code editing in an agent harness. Verify the language set, edit format and retry configuration in the original report. Benchmark documentation ↗

No primary-metric records available in this view.

· Dates may be inaccurate legacy imports; they are not verified publication dates. An unknown date is not treated as today.

How to compare a coding run

Pin the evaluation

Record the benchmark release, task window, metric, model version, reasoning effort, tool access and retry budget. A percentage from SWE-bench cannot be compared numerically with Elo from LiveCodeBench Pro.

Check your own workflow

Use representative issues with known tests. Measure successful changes, regressions, review time, latency and total cost. Preserve prompts and execution traces so another run can reproduce the result.

The previous April supplement and inferred progress charts have been withdrawn: they did not establish an October ranking or consistent evaluation protocol. Older releases remain searchable in the registry under their own names.

Agent benchmark guidance → · LLM evaluation guide → · Evidence methodology →