Browse State-of-the-Art › Reasoning
Reasoning
Benchmarks are leaderboard tables with at least one row whose task is in this area, counted by the task's area and not by the archive's per-table category tag, which tags 627 tables with Reasoning; datasets are those the archive tags with a task in this area; papers with code are catalogue papers tagged with a task in this area that list at least one repository. Task images are not shown (the archive's image host no longer serves them).
Parent tasks
26 tasks in Reasoning have sub-tasks in the archive's task tree, most benchmarks first, then most papers with code. Each section shows up to 5 sub-tasks; the task page lists them all.
- Question Answering (19)
- Classification (24)
- Natural Language Inference (3)
- Video Question Answering (2)
- Code Generation (5)
- Common Sense Reasoning (17)
- Visual Reasoning (1)
- Multi-Label Classification (4)
- Mathematical Reasoning (7)
- Logical Reasoning (22)
- 2D Human Pose Estimation (6)
- General Reinforcement Learning (2)
- Reconstruction (4)
- Program Repair (6)
- Program Synthesis (5)
- Multimodal Reasoning (1)
- Decision Making (1)
- Robot Task Planning (1)
- Mathematical Question Answering (1)
- Multi-Label Learning (1)
- Video-based Generative Performance Benchmarking (5)
- Generative Visual Question Answering (1)
- 1 Image, 2*2 Stitchi (12)
- Autonomous Navigation (3)
- Decision Making Under Uncertainty (1)
- Mathematical Proofs (1)
Question Answering
142 benchmarks · 4,171 papers with codeMultiple Choice Question Answering (MCQA)
31 benchmarks · 37 papers with code
Zero-Shot Video Question Answer
17 benchmarks · 73 papers with code
Open-Domain Question Answering
15 benchmarks · 238 papers with code
Knowledge Base Question Answering
10 benchmarks · 67 papers with code
Answer Selection
6 benchmarks · 52 papers with code
5 shown of 19 sub-tasks (3 filed under another area). All sub-tasks of Question Answering →
Classification
58 benchmarks · 3,778 papers with codeGraph Classification
73 benchmarks · 483 papers with code
Text Classification
68 benchmarks · 1,308 papers with code
Audio Classification
22 benchmarks · 183 papers with code
Medical Image Classification
11 benchmarks · 183 papers with code
Multi-class Classification
5 benchmarks · 289 papers with code
5 shown of 24 sub-tasks (6 filed under another area). All sub-tasks of Classification →
Natural Language Inference
33 benchmarks · 821 papers with codeCross-Lingual Natural Language Inference
4 benchmarks · 17 papers with code
Visual Entailment
3 benchmarks · 33 papers with code
Answer Generation
2 benchmarks · 111 papers with code
3 shown of 3 sub-tasks.
Video Question Answering
28 benchmarks · 250 papers with codeZero-Shot Video Question Answer
17 benchmarks · 73 papers with code
Few-shot Video Question Answering
0 benchmarks · 1 paper with code
2 shown of 2 sub-tasks.
Code Generation
27 benchmarks · 745 papers with codeCode Documentation Generation
7 benchmarks · 7 papers with code
Code Translation
2 benchmarks · 54 papers with code
Class-level Code Generation
1 benchmark · 4 papers with code
GitHub issue resolution
0 benchmarks · 6 papers with code
Library-Oriented Code Generation
0 benchmarks · 4 papers with code
5 shown of 5 sub-tasks (1 filed under another area).
Common Sense Reasoning
24 benchmarks · 325 papers with codeRiddle Sense
2 benchmarks · 5 papers with code
Multiview Contextual Commonsense Inference
2 benchmarks · 1 paper with code
Physical Commonsense Reasoning
1 benchmark · 6 papers with code
Discourse Marker Prediction
1 benchmark · 3 papers with code
Empirical Judgments
1 benchmark · 3 papers with code
5 shown of 17 sub-tasks (2 filed under another area). All sub-tasks of Common Sense Reasoning →
Visual Reasoning
12 benchmarks · 356 papers with codeVisual Commonsense Reasoning
7 benchmarks · 33 papers with code
1 shown of 1 sub-task.
Multi-Label Classification
10 benchmarks · 459 papers with codeHierarchical Multi-label Classification
20 benchmarks · 19 papers with code
Medical Code Prediction
7 benchmarks · 16 papers with code
Missing Labels
0 benchmarks · 50 papers with code
Extreme Multi-Label Classification
0 benchmarks · 31 papers with code
4 shown of 4 sub-tasks (2 filed under another area).
Mathematical Reasoning
10 benchmarks · 395 papers with codeMath Word Problem Solving
13 benchmarks · 80 papers with code
Formal Logic
1 benchmark · 20 papers with code
Abstract Algebra
1 benchmark · 6 papers with code
Mathematical Induction
1 benchmark · 4 papers with code
High School Mathematics
1 benchmark · 1 paper with code
5 shown of 7 sub-tasks. All sub-tasks of Mathematical Reasoning →
Logical Reasoning
10 benchmarks · 330 papers with codeTemporal Sequences
1 benchmark · 75 papers with code
Physical Intuition
1 benchmark · 15 papers with code
Elementary Mathematics
1 benchmark · 5 papers with code
Epistemic Reasoning
1 benchmark · 5 papers with code
Metaphor Boolean
1 benchmark · 3 papers with code
5 shown of 22 sub-tasks (1 filed under another area). All sub-tasks of Logical Reasoning →
2D Human Pose Estimation
10 benchmarks · 68 papers with codeAction Anticipation
8 benchmarks · 49 papers with code
Style Transfer
3 benchmarks · 759 papers with code
3D Face Animation
3 benchmarks · 25 papers with code
Community Question Answering
2 benchmarks · 50 papers with code
Semi-Supervised Human Pose Estimation
2 benchmarks · 3 papers with code
5 shown of 6 sub-tasks (3 filed under another area). All sub-tasks of 2D Human Pose Estimation →
General Reinforcement Learning
6 benchmarks · 40 papers with codeOffline RL
2 benchmarks · 310 papers with code
Model-based Reinforcement Learning
0 benchmarks · 234 papers with code
2 shown of 2 sub-tasks (1 filed under another area).
Reconstruction
5 benchmarks · 2 papers with code3D Human Reconstruction
10 benchmarks · 59 papers with code
Single-View 3D Reconstruction
9 benchmarks · 51 papers with code
Single-Image-Based Hdr Reconstruction
1 benchmark · 4 papers with code
4D reconstruction
0 benchmarks · 29 papers with code
4 shown of 4 sub-tasks (2 filed under another area).
Program Repair
4 benchmarks · 52 papers with codeFault localization
0 benchmarks · 23 papers with code
Variable misuse
0 benchmarks · 11 papers with code
Exception type
0 benchmarks · 2 papers with code
Function-docstring mismatch
0 benchmarks · 1 paper with code
Swapped operands
0 benchmarks · 1 paper with code
5 shown of 6 sub-tasks. All sub-tasks of Program Repair →
Program Synthesis
3 benchmarks · 179 papers with codeProgram Repair
4 benchmarks · 52 papers with code
Type prediction
3 benchmarks · 44 papers with code
Value prediction
1 benchmark · 21 papers with code
Enumerative Search
0 benchmarks · 5 papers with code
SQL Synthesis
0 benchmarks · 3 papers with code
5 shown of 5 sub-tasks (2 filed under another area).
Multimodal Reasoning
3 benchmarks · 138 papers with codeMME
0 benchmarks · 47 papers with code
1 shown of 1 sub-task.
Decision Making
2 benchmarks · 2,946 papers with codeImitation Learning
0 benchmarks · 691 papers with code
1 shown of 1 sub-task (1 filed under another area).
Robot Task Planning
2 benchmarks · 23 papers with codeTask Planning
0 benchmarks · 100 papers with code
1 shown of 1 sub-task.
Mathematical Question Answering
2 benchmarks · 9 papers with codeMath Word Problem Solving
13 benchmarks · 80 papers with code
1 shown of 1 sub-task.
Multi-Label Learning
1 benchmark · 92 papers with codeMissing Labels
0 benchmarks · 50 papers with code
1 shown of 1 sub-task.
Video-based Generative Performance Benchmarking
1 benchmark · 20 papers with codeVideo-based Generative Performance Benchmarking (Contextual Understanding)
1 benchmark · 16 papers with code
Video-based Generative Performance Benchmarking (Consistency)
1 benchmark · 15 papers with code
Video-based Generative Performance Benchmarking (Correctness of Information)
1 benchmark · 15 papers with code
Video-based Generative Performance Benchmarking (Detail Orientation))
1 benchmark · 15 papers with code
Video-based Generative Performance Benchmarking (Temporal Understanding)
1 benchmark · 15 papers with code
5 shown of 5 sub-tasks.
Generative Visual Question Answering
1 benchmark · 7 papers with codeVideo-based Generative Performance Benchmarking
1 benchmark · 20 papers with code
1 shown of 1 sub-task.
1 Image, 2*2 Stitchi
1 benchmark · 3 papers with codePose Estimation
31 benchmarks · 1,679 papers with code
Text-to-Image Generation
17 benchmarks · 546 papers with code
Image Deblurring
9 benchmarks · 167 papers with code
Virtual Try-on
9 benchmarks · 114 papers with code
Style Transfer
3 benchmarks · 759 papers with code
5 shown of 12 sub-tasks (10 filed under another area). All sub-tasks of 1 Image, 2*2 Stitchi →
Autonomous Navigation
0 benchmarks · 195 papers with codeAutonomous Flight (Dense Forest)
1 benchmark · 1 paper with code
Sequential Place Recognition
0 benchmarks · 5 papers with code
Autonomous Web Navigation
0 benchmarks · 4 papers with code
3 shown of 3 sub-tasks.
Decision Making Under Uncertainty
0 benchmarks · 57 papers with codeUncertainty Visualization
0 benchmarks · 5 papers with code
1 shown of 1 sub-task.
Mathematical Proofs
0 benchmarks · 29 papers with codeAutomated Theorem Proving
9 benchmarks · 110 papers with code
1 shown of 1 sub-task (1 filed under another area).
Tasks with no parent task
23 tasks in Reasoning sit at the top of the archive's task tree with no sub-tasks of their own, most benchmarks first, then most papers with code.
Arithmetic Reasoning
5 benchmarks · 112 papers with code
Error Understanding
2 benchmarks · 5 papers with code
Emotion Interpretation
2 benchmarks · 3 papers with code
Human Judgment Correlation
2 benchmarks · 3 papers with code
Natural Language Visual Grounding
1 benchmark · 30 papers with code
Odd One Out
1 benchmark · 12 papers with code
Image Paragraph Captioning
1 benchmark · 5 papers with code
Analogical Similarity
1 benchmark · 4 papers with code
Human Judgment Classification
1 benchmark · 2 papers with code
Identify Odd Metapor
1 benchmark · 2 papers with code
Commonsense Reasoning for RL
1 benchmark · 1 paper with code
ARC
0 benchmarks · 155 papers with code
Systematic Generalization
0 benchmarks · 77 papers with code
X-ray Classification
0 benchmarks · 27 papers with code
Discrete Choice Models
0 benchmarks · 16 papers with code
Causal Identification
0 benchmarks · 14 papers with code
Abstract Argumentation
0 benchmarks · 6 papers with code
Lightweight Deployment
0 benchmarks · 6 papers with code
Theory of Mind Modeling
0 benchmarks · 6 papers with code
Music Genre Transfer
0 benchmarks · 4 papers with code
Assortment Optimization
0 benchmarks · 3 papers with code
Pre-election ratings estimation
0 benchmarks · 1 paper with code
When should a hot water tank be replaced?
0 benchmarks · 1 paper with code
4 tasks in Reasoning are filed only under a parent task from another area and are not listed on this page; the parent's task page carries them.
Task tree and counts are the archive's, frozen 2025-07-28 archive 2025-07-28. Nothing here is re-ranked.