Codesota · Tasks · Vol. IICapability-first task ontologyEditorial review: October 7, 2026
§ 00 · Index

Every AI capability,
mapped to benchmark evidence.

Tasks are capabilities. Benchmarks are evidence. Domains, modalities, and safety properties are filters. This page groups 72 task pages into nine stable capability areas, then shows the canonical benchmark, recorded model, and trust grade for each row.

Reasoning, safety, robustness, multilingual coverage, and vertical domains are treated as cross-cutting overlays rather than competing top-level roots.

§ 01 · Counts

The register, by the numbers.

Registry view cached for 10 minutes · collection dates and protocol coverage vary
9
Capability areas
Stable top-level ontology
149
Tasks catalogued
72 with evidence rows
782
Datasets indexed
Canonical scope now labelled
11,851
Benchmark results
Result dates shown where recorded
§ 02 · Product map

Three different questions.

Tasks are taxonomy. Leaderboards are evidence. Lineages are benchmark history.

Tasks

Start with the problem

Use this page when you know the capability you care about: OCR, code generation, ASR, retrieval, VQA, detection.

Browse taxonomy →
Leaderboards

Then inspect the evidence

Use benchmark pages when you need result counts, source quality, trust badges, metric definitions, and current top rows.

Open leaderboards →
Lineages

Check if the benchmark still matters

Use lineage pages when a benchmark looks saturated, outdated, contaminated, or replaced by a harder successor.

View evolution →
§ 03 · Area map

Nine stable capability areas.

Domains, modalities, methods, and safety properties are filters. They do not compete with the top-level task ontology.

01 · 17 tasks
Language & Knowledge
02 · 17 tasks
Vision & Documents
03 · 7 tasks
Audio & Speech
04 · 4 tasks
Multimodal Media
05 · 7 tasks
Code & Software Engineering
06 · 9 tasks
Agents & Tool Use
07 · 6 tasks
Structured Data & Forecasting
08 · 2 tasks
Robotics, Control & RL
09 · 3 tasks
Science, Medicine & Industry
§ 04 · Capability area

Language & Knowledge.

Language understanding, retrieval, QA, RAG, factuality, and knowledge extraction. Reasoning appears here as a capability tag, not as a separate root.


Tasks
17
Registry verified flags
5
Results
583
Language & Knowledge · 17 tasks
Sorted by result count, then name
#TaskCanonical benchmarkRecorded modelScoreResults
01Multi-step Reasoning
Multi-step reasoning — maintaining coherent inference chains across 5+ sequential steps — is the meta-capabili…
Graduate-Level Google-Proof Q&A Diamond
date unknown
Claude 3.5 Sonnet
Result date unknown
59.4%
accuracy
161
02Mathematical Reasoning
Mathematical reasoning benchmarks — GSM8K, MATH, Minerva, and the competition-level AIME/AMC tests — have beco…
Mathematics Aptitude Test of Heuristics
date unknown
Apertus-70B-Instruct
Result date unknown
30.8%
accuracy
127
03Commonsense Reasoning
Commonsense reasoning — answering questions that require everyday knowledge about how the physical and social…
Massive Multitask Language Understanding
legacylegacyambiguous
MMLU is saturated and better treated as general knowledge / legacy LLM eval, not canonical commonsense reasoning.
Claude Opus 4.6
Registry date: 2026-03-01 (unverified)
91.2%
accuracy
109
04Question Answering
Question answering now spans extractive reading comprehension, open-domain retrieval QA, multi-hop reasoning,…
Natural Questions: a Benchmark for Question Answering Research
date unknown
——67
05Text Embeddings
Generating dense vector embeddings for retrieval, ranking, clustering, and semantic search.
Legacy MTEB English, 2024 snapshot
historicallegacyambiguous
NV-Embed-v2 is a historical MTEB English 56-task snapshot, not a fresh 2026 embedding frontier.
——44
06Text Summarization
Text summarization compresses documents while preserving key information — a task that became dramatically mor…
CNN/DailyMail Summarization
stale
Llama 3.1 405B
Registry date: 2024-07-31 (unverified)
45.1%
rouge-1
16
07Logical Reasoning
Logical reasoning — formal deduction, constraint satisfaction, and syllogistic inference — exposes a core weak…
LogiQA
date unknown
Claude 3.5 Sonnet
Result date unknown
53.8%
accuracy
12
08Text Ranking
Text ranking is the invisible backbone of every search engine and RAG pipeline. The field was transformed by C…
BEIR legacy retrieval
legacylegacyambiguous
Legacy retrieval snapshot. Split modern retrieval, reranking, multilingual, and long-context RAG evals before calling this current SOTA.
——9
09Natural Language Inference
Determining entailment relationships between sentences (SNLI, MNLI).
Stanford Natural Language Inference
stale
Llama 3.1 405B
Registry date: 2024-07-31 (unverified)
91.2%
accuracy
8
10Named Entity Recognition
Named entity recognition (NER) extracts structured mentions — people, organizations, locations, dates — from u…
CoNLL-2003 Named Entity Recognition
stale
Llama 3.1 405B
Registry date: 2024-07-31 (unverified)
90.6%
f1
7
11Arithmetic Reasoning
Arithmetic reasoning — solving computation-heavy problems stated in natural language — tests whether models ca…
Math Word Problem Repository
date unknown
Claude 3.5 Sonnet
Result date unknown
95.8%
accuracy
6
12Entity Linking
Linking mentions to knowledge base entities.
AIDA-CoNLL-YAGO (test-b)
date unknown
GENRE
Result date unknown
93.30
micro_f1
3
13Knowledge Graph Completion
Predicting missing links in knowledge graphs.
FB15k-237 Knowledge Graph Completion
date unknown
NBFNet
Result date unknown
0.415
mrr
3
14Relation Extraction
Extracting relationships between entities from text.
TAC Relation Extraction Dataset
date unknown
KEPLER
Result date unknown
71.7%
f1
3
15Semantic Textual Similarity
Semantic similarity measures how close two pieces of text are in meaning — the foundation of duplicate detecti…
STS Benchmark
stale
E5-Mistral-7B-instruct
Registry date: 2024-01-01 (unverified)
84.70
spearman
3
16Table Question Answering
Table question answering bridges natural language and structured data — asking "what was Q3 revenue?" over a s…
WikiTableQuestions
stale
TAPAS-large
Registry date: 2020-04-06 (unverified)
48.7%
accuracy
3
17Reading Comprehension
Understanding and answering questions about passages.
ReAding Comprehension from Examinations
date unknown
ALBERT ensemble
Result date unknown
89.4%
accuracy
2
Fig 04 · Each row links to the task page with full history. Shaded rows mark independently verified state of the art; empty score cells mean no benchmark in the register yet clears our trust bar.
§ 05 · Capability area

Vision & Documents.

Images, video frames, OCR, layout, tables, document parsing, detection, segmentation, and visual anomaly detection.


Tasks
17
Registry verified flags
2
Results
2,326
Vision & Documents · 17 tasks
Sorted by result count, then name
#TaskCanonical benchmarkRecorded modelScoreResults
01Document OCR
Reading text, structure, and layout from document images.
OCRBench v2 public overall
submetricdate unknownambiguous
Scope is public overall. Do not compare directly with English-private OCRBench v2 or full document parsing metrics.
——831
02Scene Text Detection
Detecting text regions in natural scene images
COCO-Text detection scope needs review
misclassifieddate unknownmisclassified
CLIP4STR-style scene text recognition rows do not belong under detection. Detection needs region metrics such as precision, recall, F-measure, or hmean.
——581
03Document Parsing
Parsing document structure and content
OmniDocBench v1.5
submetricdate unknownambiguous
Reading order is only one OmniDocBench facet. Summary SOTA needs text, layout, table TEDS, reading order, and end-to-end structure facets.
clearOCR
Result date unknown
31.70
composite
149
04Document Layout Analysis
Analyzing the layout structure of documents
d4la
date unknown
——133
05Scene Text Recognition
Recognizing text in natural scene images
cute80
stale
DTrOCR 105M
Registry date: 2023-08-30 (unverified)
99.1%
accuracy
127
06Object Detection
Object Detection is a computer vision task that involves identifying and localizing objects within an image. T…
Microsoft Common Objects in Context
stale
ScyllaNet
Registry date: 2025-09-01 (unverified)
66.12
box-map
104
07Image Classification
Image Classification is a fundamental task in computer vision that aims to assign a label or class to an entir…
ImageNet Large Scale Visual Recognition Challenge 2012
date unknown
AIMv2-3B
Result date unknown
89.50
top-1-accuracy
87
08Table Recognition
Detecting and parsing tables in documents
ICDAR2013 table structure (legacy)
legacylegacyambiguous
ICDAR2013 is too narrow for 2026 table recognition. Promote PubTables-1M, PubTabNet, FinTabNet, or table-specific document parsing metrics.
——71
09General OCR Capabilities
Comprehensive benchmarks covering multiple aspects of OCR performance.
OCRBench v2
needs coveragedate unknownambiguous
Fold this into OCR unless the metric scope is explicit: public overall, English-private, recognition, understanding, or full parsing.
——70
10Document Image Classification
Classifying documents by type or category
aip
date unknown
——63
11Handwriting Recognition
Recognizing handwritten text
—
date unknown
——40
12Document Understanding
Document understanding requires parsing visually rich documents — invoices, forms, scientific papers, tables —…
Form Understanding in Noisy Scanned Documents
date unknown
——28
13Semantic Segmentation
Semantic segmentation assigns a class label to every pixel — the dense prediction problem that underpins auton…
ADE20K Scene Parsing Benchmark
date unknown
BEiT-3 (ViT-L)
Result date unknown
62.8%
mIoU
24
14Video classification
The task of classifying videos into predefined categories or classes. Video classification involves analyzing…
Kinetics-400
date unknown
——13
15Image segmentation
Image segmentation is a computer vision technique that divides a digital image into multiple parts or "segment…
—
date unknown
——3
16Keypoint Detection
Keypoint detection localizes specific anatomical or structural landmarks — body joints, facial features, hand…
COCO Keypoints
date unknown
ViTPose-G
Result date unknown
80.9%
map
1
17OCR
OCR, or Optical Character Recognition, is the task of converting an image containing text into machine-readabl…
—
date unknown
——1
Fig 05 · Each row links to the task page with full history. Shaded rows mark independently verified state of the art; empty score cells mean no benchmark in the register yet clears our trust bar.
§ 06 · Capability area

Audio & Speech.

ASR, TTS, speaker intelligence, music, sound events, audio-language understanding, and audio safety.


Tasks
7
Registry verified flags
1
Results
545
Audio & Speech · 7 tasks
Sorted by result count, then name
#TaskCanonical benchmarkRecorded modelScoreResults
01Speech Recognition
Automatic speech recognition went from a specialized pipeline (acoustic model + language model + decoder) to a…
Mozilla Common Voice
stale
Whisper Large v2
Registry date: 2022-12-06 (unverified)
11.20
wer
526
02Audio Captioning
Generating text descriptions of audio content.
AudioCaps
historicaldate unknownambiguous
Baseline-style AudioCaps rows should not read as current leading audio-language SOTA without a refresh.
AudioCaps baseline (TopDown+Align)
Result date unknown
36.9%
spider
7
03Music Generation
Generating music from text, audio, or other inputs.
MusicCaps
historicaldate unknownambiguous
MusicLM is historically important, but this needs MusicCaps/MusicBench, human eval, and proprietary/open splits.
MusicGen Large
Result date unknown
3.800
fad
3
04Sound Event Detection
Detecting and localizing sound events in audio.
Domestic Environment Sound Event Detection (DCASE Task 4)
date unknown
ATST-SED
Result date unknown
58.10
event-f1
3
05Speaker Verification
Verifying speaker identity from voice samples.
VoxCeleb1 Original Test Set (VoxCeleb1-O)
date unknown
ECAPA-TDNN
Result date unknown
0.870
eer
3
06Speech Translation
Translating spoken audio directly to another language.
MuST-C English-German tst-COMMON
date unknown
Fairseq S2T (MuST-C)
Result date unknown
22.7%
bleu
3
07Speech Enhancement
Recovering clean speech from noisy recordings. Benchmarked on VoiceBank+DEMAND (PESQ, STOI, SI-SDR) and the Mi…
—
date unknown
——0
Fig 06 · Each row links to the task page with full history. Shaded rows mark independently verified state of the art; empty score cells mean no benchmark in the register yet clears our trust bar.
§ 07 · Capability area

Multimodal Media.

Cross-modal tasks only: VQA, image-text retrieval, video QA, document VQA, text-to-image, image editing, and any-to-any media models.


Tasks
4
Registry verified flags
1
Results
206
Multimodal Media · 4 tasks
Sorted by result count, then name
#TaskCanonical benchmarkRecorded modelScoreResults
01Visual Question Answering
Visual question answering (VQA) is the original multimodal reasoning task — given an image and a natural langu…
Visual Question Answering v2.0
stale
GPT-4o
Registry date: 2024-10-25 (unverified)
78.5%
accuracy
147
02Video Understanding
Video understanding asks models to reason over temporal sequences — answering questions, generating summaries,…
MVBench
date unknown
LongCat-Flash-Omni
Result date unknown
75.2%
accuracy
44
03Text-to-Image Generation
Text-to-image generation went from "interesting research" to cultural phenomenon in 18 months. DALL-E 2 (2022)…
DPG-Bench
date unknown
——8
04Image Captioning
Image captioning — generating natural language descriptions of images — was the task that launched the modern…
COCO Captions
legacylegacyambiguous
COCO captioning is legacy and saturated. Add NoCaps, Flickr30k, caption QA, or preference-based caption evals.
Chameleon-SFT
Result date unknown
140.8%
cider
7
Fig 07 · Each row links to the task page with full history. Shaded rows mark independently verified state of the art; empty score cells mean no benchmark in the register yet clears our trust bar.
§ 08 · Capability area

Code & Software Engineering.

Code generation, completion, repair, repository understanding, tests, vulnerability work, UI code, and mobile app code generation.


Tasks
7
Registry verified flags
5
Results
337
Code & Software Engineering · 7 tasks
Sorted by result count, then name
#TaskCanonical benchmarkRecorded modelScoreResults
01Code Generation
Generating code from natural language descriptions (HumanEval, MBPP).
LiveCodeBench
date unknown
——270
02React Native Code Generation
Evaluating AI models on generating correct, production-quality React Native implementations. Covers animation,…
Callstack Incubator React Native Evaluation Suite
date unknown
Claude Opus 4.6
Result date unknown
84.36
requirement-satisfaction
40
03Code Translation
Converting code between programming languages.
TransCoder Evaluation on GeeksForGeeks Algorithmic Problems
stale
Qwen2.5-Coder 32B
Registry date: 2024-09-19 (unverified)
86.30
computational-accuracy
7
04Bug Detection
Identifying bugs and vulnerabilities in code.
Bugs2Fix: Learning to Rewrite Buggy Code
stale
Qwen2.5-Coder 32B
Registry date: 2024-09-19 (unverified)
76.8%
accuracy
6
05Code Completion
Predicting the next tokens in code sequences.
Cross-File Code Completion Evaluation
stale
Qwen2.5-Coder 32B
Registry date: 2024-09-19 (unverified)
43.70
exact-match
6
06Program Repair
Automatically fixing bugs in code.
Defects4J: A Database of Real Faults in Java Programs
stale
GPT-4o
Registry date: 2024-04-18 (unverified)
82.00
correct-patches
5
07Code Summarization
Generating natural language descriptions of code.
CodeXGLUE Code-to-Text Python subset
date unknown
CodeBERT
Result date unknown
19.1%
bleu
3
Fig 08 · Each row links to the task page with full history. Shaded rows mark independently verified state of the art; empty score cells mean no benchmark in the register yet clears our trust bar.
§ 09 · Capability area

Agents & Tool Use.

Tool calling, web and desktop agents, browser automation, long-horizon autonomy, multi-agent coordination, and agent safety.


Tasks
9
Registry verified flags
6
Results
225
Agents & Tool Use · 9 tasks
Sorted by result count, then name
#TaskCanonical benchmarkRecorded modelScoreResults
01SWE-bench
SWE-bench — resolving real GitHub issues from popular Python repositories — became the defining benchmark for…
SWE-bench Verified — Agentic Leaderboard
date unknown
Claude 3.5 Haiku
Result date unknown
40.60
resolve-rate
81
02Task agents
AI agents are autonomous software systems that use artificial intelligence to achieve goals and complete tasks…
Collider-Bench simulation tasks
date unknown
Claude Code (Haiku 4.5)
Result date unknown
0.000
acc-tau-0-33
45
03Web & Desktop Agents
Web and desktop agents — AI systems that operate browsers and GUIs to complete real tasks — are benchmarked by…
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
stale
Claude Computer Use
Registry date: 2024-04-11 (unverified)
14.90
success-rate
39
04Autonomous Coding
Agent benchmarks where systems complete coding, terminal, repository, or developer-workflow tasks with minimal…
SWE-bench Verified (Agentic)
date unknown
Claude Opus 4.5
Result date unknown
80.90
pct_resolved
23
05Tool Use
Benchmarks measuring AI agents ability to use tools and APIs to complete real-world tasks across domains like…
—
date unknown
——19
06HCAST
HCAST (Human-Calibrated Autonomy Software Tasks) is a 90-task benchmark from METR designed to measure AI auton…
Human-Calibrated Autonomy Software Tasks
stale
Claude 3.5 Sonnet
Registry date: 2025-04-01 (unverified)
18.00
success-rate
6
07RE-Bench
RE-Bench (Research Engineering Benchmark) from METR evaluates AI agents on 7 open-ended ML research engineerin…
Research Engineering Benchmark
stale
Claude 3.7 Sonnet
Registry date: 2025-04-01 (unverified)
0.290
normalized-score
5
08Time Horizon
Time horizon — how long an AI agent can work autonomously before requiring human correction — is arguably the…
METR Autonomy Evaluation: Time Horizon
stale
Claude 3.7 Sonnet
Registry date: 2025-04-01 (unverified)
14.00
task-horizon-minutes
5
09Bioinformatics Agents
LLM-agent benchmarks for computational biology — exploring datasets, running multi-step analyses, and interpre…
—
date unknown
——2
Fig 09 · Each row links to the task page with full history. Shaded rows mark independently verified state of the art; empty score cells mean no benchmark in the register yet clears our trust bar.
§ 10 · Capability area

Structured Data & Forecasting.

Tables, tabular classification and regression, time-series forecasting, anomaly detection, recommender systems, graph learning, and optimization.


Tasks
6
Registry verified flags
3
Results
19
Structured Data & Forecasting · 6 tasks
Sorted by result count, then name
#TaskCanonical benchmarkRecorded modelScoreResults
01Node Classification
Node classification — assigning labels to vertices in a graph using both node features and neighborhood struct…
Cora Citation Network
stale
ACNet
Registry date: 2019-04-07 (unverified)
83.5%
accuracy
6
02Tabular Classification
Tabular classification — predicting discrete labels from structured rows and columns — remains the one domain…
OpenML-CC18
stale
AutoGluon-Tabular
Registry date: 2025-06-01 (unverified)
88.5%
accuracy
5
03Link Prediction
Link prediction — inferring missing or future edges in a graph — underpins knowledge graph completion, drug-ta…
Open Graph Benchmark - ogbl-collab
date unknown
BUDDY
Result date unknown
65.94
hits_at_50
3
04Molecular Property Prediction
Molecular property prediction — estimating toxicity, solubility, binding affinity, or other properties from mo…
Open Graph Benchmark - ogbg-molhiv
date unknown
DGN
Result date unknown
79.70
roc_auc
3
05Tabular Regression
Tabular regression — predicting continuous values from structured data — powers everything from house-price es…
California Housing
date unknown
LightGBM
Result date unknown
0.433
rmse
2
06Tabular Machine Learning
Classification and regression on tabular data. Benchmarked on TabArena (Elo over 51 datasets); gradient-booste…
—
date unknown
——0
Fig 10 · Each row links to the task page with full history. Shaded rows mark independently verified state of the art; empty score cells mean no benchmark in the register yet clears our trust bar.
§ 11 · Capability area

Robotics, Control & RL.

Game playing, continuous control, manipulation, navigation, embodied instruction following, VLA models, drones, and autonomous driving.


Tasks
2
Registry verified flags
0
Results
21
Robotics, Control & RL · 2 tasks
Sorted by result count, then name
#TaskCanonical benchmarkRecorded modelScoreResults
01Atari Games
Atari games became the canonical RL benchmark when DeepMind's DQN (2013) learned to play Breakout from raw pix…
Arcade Learning Environment (Atari 2600)
legacylegacy
Classic RL benchmark. Keep separate from modern embodied, VLA, robotics manipulation, and navigation tasks.
Agent57
Result date unknown
4731.3
human-normalized-score
12
02Continuous Control
Continuous control — learning smooth motor commands in simulated physics — was transformed by MuJoCo and the O…
Multi-Joint dynamics with Contact
submetricdate unknownambiguous
MuJoCo control is a narrow simulation slice. Split from robotics manipulation, navigation, and VLA evaluations.
BRO
Result date unknown
941.0
average-return
9
Fig 11 · Each row links to the task page with full history. Shaded rows mark independently verified state of the art; empty score cells mean no benchmark in the register yet clears our trust bar.
§ 12 · Capability area

Science, Medicine & Industry.

A domain layer for medical imaging, clinical text, drug discovery, protein modeling, industrial inspection, remote sensing, climate, legal, finance, and compliance AI.


Tasks
3
Registry verified flags
3
Results
110
Science, Medicine & Industry · 3 tasks
Sorted by result count, then name
#TaskCanonical benchmarkRecorded modelScoreResults
01Disease Classification
Diagnosing diseases from medical images or data.
Autism Brain Imaging Data Exchange I
claim-onlystaleambiguous
High ABIDE accuracy claims are leakage-risk until subject-level split, site-held-out validation, preprocessing, confound control, and external validation are verified.
ChebGAT-GCN
Registry date: 2025-11-27 (unverified)
74.8%
accuracy
57
02Anomaly Detection
Detecting defects and anomalies in manufacturing (MVTec AD, VisA).
MVTec Anomaly Detection Dataset
submetricstaleambiguous
MVTec AD rows must split image-level classification, pixel-level localization, zero/few/full-shot, and AUROC/AUPRO metric scopes.
AnomalyGPT
Registry date: 2023-08-29 (unverified)
97.40
auroc
27
03Medical Image Segmentation
Segmenting organs and abnormalities in medical images.
Automated Cardiac Diagnosis Challenge
stale
LightM-UNet
Registry date: 2024-03-04 (unverified)
90.49
mean-dsc
26
Fig 12 · Each row links to the task page with full history. Shaded rows mark independently verified state of the art; empty score cells mean no benchmark in the register yet clears our trust bar.
§ 13
Trust grades

What the letters mean.

Benchmarks are not equally believable. Some are held out behind a private evaluator; some ship their test set as part of the training corpus. We grade the canonical dataset of every task on a four-point scale and show the letter next to the score.

A
Reproduced · dated · code
The full path is visible: a public checkpoint, a frozen commit, a declared environment, and a score we (or a signed reproducer) ran against a held-out test set. Contamination controlled, metric direction declared, date stamped.
B
Partial reproduction
Known weaknesses — evaluator overlap, public answer keys, a missing seed — but the submission otherwise checks out. Cite with caution; we preserve the caveat alongside the number.
C
Claim-only
The authors say so. We have not reproduced it and cannot yet. Shown in the register for completeness, but do not treat as state of the art.
F
Contested or retracted
The benchmark is considered unreliable: documented contamination, split leakage, or a score withdrawn by its authors. The row remains visible — leaderboards that silently forget are worse than leaderboards that argue in public.

A dataset can be regraded in public at any time; the history is preserved on the benchmark page. We publish the regrade, we don't erase the prior.

§ 14 · Standing columns

Capability buckets, not benchmarks.

HuggingFace pipeline-tag categories. These group concrete tasks thematically; they are not themselves measurable. Use them to navigate to the real rankings.

Standing column

Image + Text → Video

Animate a still image guided by a text prompt.

Standing column

Video → Video

Video editing, style transfer, super-resolution.

Standing column

Image → 3D

Generate a 3D mesh or NeRF from one or more images.

Standing column

Text → 3D

Generate a 3D asset from a text prompt.

Standing column

Image → Video

Animate a still image into a short clip.

Standing column

Unconditional Image Generation

Generative image models without text conditioning (DCGAN, StyleGAN era).

Fig 14 · Standing columns exist to aid navigation, not to be ranked. Follow any link to the underlying task's leaderboard.
§ 15
Methodology

Why this register can be trusted.

Most leaderboards are a ledger of claims. Authors submit a number, a banner appears; the number stands until the next banner appears. Codesota is different in three ordinary ways.

First, every submission carries code. Not a repo link alone — a frozen commit, a declared environment, a recorded seed. If it does not run, the row does not publish.

Second, every benchmark has a metric direction. Higher-is-better and lower-is-better are declared on the dataset; no ambiguity reaches the reader.

Third, every score carries a date. When a model regresses — and they do — the record is preserved. The table never silently forgets.