Stanford x Laude benchmark for AI agents operating in terminal environments. Terminal-Bench 2.0 evaluates terminal mastery across software engineering, machine learning, security, data science, system administration, file operations, and related operational workflows. Official site lists 89 high-quality tasks and a 124-entry live leaderboard.
Accuracy is the reported evaluation metric for Terminal-Bench 2.0. Codesota tracks published model scores on this metric so readers can compare state-of-the-art results across sources and model families.
Higher is better
Muted rows were not state of the art when published — an earlier or same-year result already scored better.