7,787 science questions requiring reasoning. Challenge set contains harder questions that retrieval fails on.
Accuracy is the reported evaluation metric for ARC-Challenge. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.
| Rank | Model | Trust | Score | Year | Links | Fix |
|---|---|---|---|---|---|---|
| 01 | claude-35-sonnet | unverified | 96.7 | 2026 | N/A | Looks wrong? |
| 02 | gpt-4o | unverified | 96.4 | 2026 | N/A | Looks wrong? |
| 03 | gemini-15-pro | unverified | 94.8 | 2026 | N/A | Looks wrong? |
| 04 | llama-3-70b | unverified | 93 | 2026 | N/A | Looks wrong? |
Questions, claims, and experiments linked through CodeSOTA’s public evidence graph.