Recent Papers / arXiv:2605.13950

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction

arXiv:2605.13950Submitted Jan 1, 20266 benchmark results

Darius A. Faroughy, Sofia Palacios Schweitzer, Ian Pang, Siddharth Mishra-Sharma, David Shih

Abstract

Collider-Bench is a benchmark for evaluating autonomous LLM agents on long-horizon, real-world scientific tasks involving the reproduction of Large Hadron Collider (LHC) experimental analyses.

Agents must turn published papers into executable simulation-and-selection pipelines to predict collision event yields, evaluated against quantitative targets.

Tasks
edit
Results

6 results reproduced from this paper.

submit
Sorted instantly in-page
Results
6
SOTA rows
1
Models
6
Datasets
1
CodeSOTA extraction

Benchmark evidence

edit

Link this paper to benchmark rows, datasets, model cards, and reproduced results as evidence is extracted.

§ 02 · Models

6 models from this paper.

evaluates
Claude Code (Haiku 4.5)
Anthropic
evaluates
Claude Code (Opus 4.7)
Anthropic
evaluates
Claude Code (Sonnet 4.6)
Anthropic
evaluates
Codex CLI (GPT-5.4-mini)
OpenAI
evaluates
Codex CLI (GPT-5.5)
OpenAI
evaluates
ForgeCode (DeepSeek-V4)
DeepSeek
§ 03 · Datasets

1 dataset from this paper.

uses · Agentic AI
Collider-Bench
Task agents
Add or update benchmark results
Logged-in editor · benchmark trail
Log in is not configured in this environment.
Read next

Three places to go from here.

Index
All papers
All tracked papers in the registry, with benchmark result, model, and leaderboard linkage where available.
Replacement
Papers with Code is dead — alternatives
What replaced PWC for each use case: LLMs, OCR, speech, vision, robotics.
Top hub
LLM benchmarks
Every frontier LLM benchmark, scored.