| 01 | Multi-step Reasoning Multi-step reasoning — maintaining coherent inference chains across 5+ sequential steps — is the meta-capabili… | Graduate-Level Google-Proof Q&A Diamond date unknown | Claude 3.5 Sonnet Result date unknown | 59.4% accuracy | 161 |
| 02 | Mathematical Reasoning Mathematical reasoning benchmarks — GSM8K, MATH, Minerva, and the competition-level AIME/AMC tests — have beco… | Mathematics Aptitude Test of Heuristics date unknown | Apertus-70B-Instruct Result date unknown | 30.8% accuracy | 127 |
| 03 | Commonsense Reasoning Commonsense reasoning — answering questions that require everyday knowledge about how the physical and social… | Massive Multitask Language Understanding legacylegacyambiguous MMLU is saturated and better treated as general knowledge / legacy LLM eval, not canonical commonsense reasoning. | Claude Opus 4.6 Registry date: 2026-03-01 (unverified) | 91.2% accuracy | 109 |
| 04 | Question Answering Question answering now spans extractive reading comprehension, open-domain retrieval QA, multi-hop reasoning,… | Natural Questions: a Benchmark for Question Answering Research date unknown | — | — | 67 |
| 05 | Text Embeddings Generating dense vector embeddings for retrieval, ranking, clustering, and semantic search. | Legacy MTEB English, 2024 snapshot historicallegacyambiguous NV-Embed-v2 is a historical MTEB English 56-task snapshot, not a fresh 2026 embedding frontier. | — | — | 44 |
| 06 | Text Summarization Text summarization compresses documents while preserving key information — a task that became dramatically mor… | CNN/DailyMail Summarization stale | Llama 3.1 405B Registry date: 2024-07-31 (unverified) | 45.1% rouge-1 | 16 |
| 07 | Logical Reasoning Logical reasoning — formal deduction, constraint satisfaction, and syllogistic inference — exposes a core weak… | LogiQA date unknown | Claude 3.5 Sonnet Result date unknown | 53.8% accuracy | 12 |
| 08 | Text Ranking Text ranking is the invisible backbone of every search engine and RAG pipeline. The field was transformed by C… | BEIR legacy retrieval legacylegacyambiguous Legacy retrieval snapshot. Split modern retrieval, reranking, multilingual, and long-context RAG evals before calling this current SOTA. | — | — | 9 |
| 09 | Natural Language Inference Determining entailment relationships between sentences (SNLI, MNLI). | Stanford Natural Language Inference stale | Llama 3.1 405B Registry date: 2024-07-31 (unverified) | 91.2% accuracy | 8 |
| 10 | Named Entity Recognition Named entity recognition (NER) extracts structured mentions — people, organizations, locations, dates — from u… | CoNLL-2003 Named Entity Recognition stale | Llama 3.1 405B Registry date: 2024-07-31 (unverified) | 90.6% f1 | 7 |
| 11 | Arithmetic Reasoning Arithmetic reasoning — solving computation-heavy problems stated in natural language — tests whether models ca… | Math Word Problem Repository date unknown | Claude 3.5 Sonnet Result date unknown | 95.8% accuracy | 6 |
| 12 | Entity Linking Linking mentions to knowledge base entities. | AIDA-CoNLL-YAGO (test-b) date unknown | GENRE Result date unknown | 93.30 micro_f1 | 3 |
| 13 | Knowledge Graph Completion Predicting missing links in knowledge graphs. | FB15k-237 Knowledge Graph Completion date unknown | NBFNet Result date unknown | 0.415 mrr | 3 |
| 14 | Relation Extraction Extracting relationships between entities from text. | TAC Relation Extraction Dataset date unknown | KEPLER Result date unknown | 71.7% f1 | 3 |
| 15 | Semantic Textual Similarity Semantic similarity measures how close two pieces of text are in meaning — the foundation of duplicate detecti… | STS Benchmark stale | E5-Mistral-7B-instruct Registry date: 2024-01-01 (unverified) | 84.70 spearman | 3 |
| 16 | Table Question Answering Table question answering bridges natural language and structured data — asking "what was Q3 revenue?" over a s… | WikiTableQuestions stale | TAPAS-large Registry date: 2020-04-06 (unverified) | 48.7% accuracy | 3 |
| 17 | Reading Comprehension Understanding and answering questions about passages. | ReAding Comprehension from Examinations date unknown | ALBERT ensemble Result date unknown | 89.4% accuracy | 2 |