Browse State-of-the-Art › Reasoning

Reasoning

603 benchmarks 166 tasks 1,159 datasets 16,693 papers with code archive 2025-07-28

Syntology code harvested from 4,382 of the papers with code counted above at least one sample ran for 3,405 of them per-sample status is on each paper page

Benchmarks are leaderboard tables with at least one row whose task is in this area, counted by the task's area and not by the archive's per-table category tag, which tags 627 tables with Reasoning; datasets are those the archive tags with a task in this area; papers with code are catalogue papers tagged with a task in this area that list at least one repository. Task images are not shown (the archive's image host no longer serves them).

Parent tasks

26 tasks in Reasoning have sub-tasks in the archive's task tree, most benchmarks first, then most papers with code. Each section shows up to 5 sub-tasks; the task page lists them all.

Question Answering

142 benchmarks · 4,171 papers with code

Multiple Choice Question Answering (MCQA)

31 benchmarks · 37 papers with code

Zero-Shot Video Question Answer

17 benchmarks · 73 papers with code

Open-Domain Question Answering

15 benchmarks · 238 papers with code

Knowledge Base Question Answering

10 benchmarks · 67 papers with code

Answer Selection

6 benchmarks · 52 papers with code

5 shown of 19 sub-tasks (3 filed under another area). All sub-tasks of Question Answering →

Classification

58 benchmarks · 3,778 papers with code

Graph Classification

73 benchmarks · 483 papers with code

Text Classification

68 benchmarks · 1,308 papers with code

Audio Classification

22 benchmarks · 183 papers with code

Medical Image Classification

11 benchmarks · 183 papers with code

Multi-class Classification

5 benchmarks · 289 papers with code

5 shown of 24 sub-tasks (6 filed under another area). All sub-tasks of Classification →

Natural Language Inference

33 benchmarks · 821 papers with code

Cross-Lingual Natural Language Inference

4 benchmarks · 17 papers with code

Visual Entailment

3 benchmarks · 33 papers with code

Answer Generation

2 benchmarks · 111 papers with code

3 shown of 3 sub-tasks.

Video Question Answering

28 benchmarks · 250 papers with code

Zero-Shot Video Question Answer

17 benchmarks · 73 papers with code

Few-shot Video Question Answering

0 benchmarks · 1 paper with code

2 shown of 2 sub-tasks.

Code Generation

27 benchmarks · 745 papers with code

Code Documentation Generation

7 benchmarks · 7 papers with code

Code Translation

2 benchmarks · 54 papers with code

Class-level Code Generation

1 benchmark · 4 papers with code

GitHub issue resolution

0 benchmarks · 6 papers with code

Library-Oriented Code Generation

0 benchmarks · 4 papers with code

5 shown of 5 sub-tasks (1 filed under another area).

Common Sense Reasoning

24 benchmarks · 325 papers with code

Riddle Sense

2 benchmarks · 5 papers with code

Multiview Contextual Commonsense Inference

2 benchmarks · 1 paper with code

Physical Commonsense Reasoning

1 benchmark · 6 papers with code

Discourse Marker Prediction

1 benchmark · 3 papers with code

Empirical Judgments

1 benchmark · 3 papers with code

5 shown of 17 sub-tasks (2 filed under another area). All sub-tasks of Common Sense Reasoning →

Visual Reasoning

12 benchmarks · 356 papers with code

Visual Commonsense Reasoning

7 benchmarks · 33 papers with code

1 shown of 1 sub-task.

Multi-Label Classification

10 benchmarks · 459 papers with code

Hierarchical Multi-label Classification

20 benchmarks · 19 papers with code

Medical Code Prediction

7 benchmarks · 16 papers with code

Missing Labels

0 benchmarks · 50 papers with code

Extreme Multi-Label Classification

0 benchmarks · 31 papers with code

4 shown of 4 sub-tasks (2 filed under another area).

Mathematical Reasoning

10 benchmarks · 395 papers with code

Math Word Problem Solving

13 benchmarks · 80 papers with code

Formal Logic

1 benchmark · 20 papers with code

Abstract Algebra

1 benchmark · 6 papers with code

Mathematical Induction

1 benchmark · 4 papers with code

High School Mathematics

1 benchmark · 1 paper with code

5 shown of 7 sub-tasks. All sub-tasks of Mathematical Reasoning →

Logical Reasoning

10 benchmarks · 330 papers with code

Temporal Sequences

1 benchmark · 75 papers with code

Physical Intuition

1 benchmark · 15 papers with code

Elementary Mathematics

1 benchmark · 5 papers with code

Epistemic Reasoning

1 benchmark · 5 papers with code

Metaphor Boolean

1 benchmark · 3 papers with code

5 shown of 22 sub-tasks (1 filed under another area). All sub-tasks of Logical Reasoning →

2D Human Pose Estimation

10 benchmarks · 68 papers with code

Action Anticipation

8 benchmarks · 49 papers with code

Style Transfer

3 benchmarks · 759 papers with code

3D Face Animation

3 benchmarks · 25 papers with code

Community Question Answering

2 benchmarks · 50 papers with code

Semi-Supervised Human Pose Estimation

2 benchmarks · 3 papers with code

5 shown of 6 sub-tasks (3 filed under another area). All sub-tasks of 2D Human Pose Estimation →

General Reinforcement Learning

6 benchmarks · 40 papers with code

Offline RL

2 benchmarks · 310 papers with code

Model-based Reinforcement Learning

0 benchmarks · 234 papers with code

2 shown of 2 sub-tasks (1 filed under another area).

Reconstruction

5 benchmarks · 2 papers with code

3D Human Reconstruction

10 benchmarks · 59 papers with code

Single-View 3D Reconstruction

9 benchmarks · 51 papers with code

Single-Image-Based Hdr Reconstruction

1 benchmark · 4 papers with code

4D reconstruction

0 benchmarks · 29 papers with code

4 shown of 4 sub-tasks (2 filed under another area).

Program Repair

4 benchmarks · 52 papers with code

Fault localization

0 benchmarks · 23 papers with code

Variable misuse

0 benchmarks · 11 papers with code

Exception type

0 benchmarks · 2 papers with code

Function-docstring mismatch

0 benchmarks · 1 paper with code

Swapped operands

0 benchmarks · 1 paper with code

5 shown of 6 sub-tasks. All sub-tasks of Program Repair →

Program Synthesis

3 benchmarks · 179 papers with code

Program Repair

4 benchmarks · 52 papers with code

Type prediction

3 benchmarks · 44 papers with code

Value prediction

1 benchmark · 21 papers with code

Enumerative Search

0 benchmarks · 5 papers with code

SQL Synthesis

0 benchmarks · 3 papers with code

5 shown of 5 sub-tasks (2 filed under another area).

Multimodal Reasoning

3 benchmarks · 138 papers with code

MME

0 benchmarks · 47 papers with code

1 shown of 1 sub-task.

Decision Making

2 benchmarks · 2,946 papers with code

Imitation Learning

0 benchmarks · 691 papers with code

1 shown of 1 sub-task (1 filed under another area).

Robot Task Planning

2 benchmarks · 23 papers with code

Task Planning

0 benchmarks · 100 papers with code

1 shown of 1 sub-task.

Mathematical Question Answering

2 benchmarks · 9 papers with code

Math Word Problem Solving

13 benchmarks · 80 papers with code

1 shown of 1 sub-task.

Multi-Label Learning

1 benchmark · 92 papers with code

Missing Labels

0 benchmarks · 50 papers with code

1 shown of 1 sub-task.

Video-based Generative Performance Benchmarking

1 benchmark · 20 papers with code

5 shown of 5 sub-tasks.

Generative Visual Question Answering

1 benchmark · 7 papers with code

Video-based Generative Performance Benchmarking

1 benchmark · 20 papers with code

1 shown of 1 sub-task.

1 Image, 2*2 Stitchi

1 benchmark · 3 papers with code

Pose Estimation

31 benchmarks · 1,679 papers with code

Text-to-Image Generation

17 benchmarks · 546 papers with code

Image Deblurring

9 benchmarks · 167 papers with code

Virtual Try-on

9 benchmarks · 114 papers with code

Style Transfer

3 benchmarks · 759 papers with code

5 shown of 12 sub-tasks (10 filed under another area). All sub-tasks of 1 Image, 2*2 Stitchi →

Autonomous Navigation

0 benchmarks · 195 papers with code

Autonomous Flight (Dense Forest)

1 benchmark · 1 paper with code

Sequential Place Recognition

0 benchmarks · 5 papers with code

Autonomous Web Navigation

0 benchmarks · 4 papers with code

3 shown of 3 sub-tasks.

Decision Making Under Uncertainty

0 benchmarks · 57 papers with code

Uncertainty Visualization

0 benchmarks · 5 papers with code

1 shown of 1 sub-task.

Mathematical Proofs

0 benchmarks · 29 papers with code

Automated Theorem Proving

9 benchmarks · 110 papers with code

1 shown of 1 sub-task (1 filed under another area).

Tasks with no parent task

23 tasks in Reasoning sit at the top of the archive's task tree with no sub-tasks of their own, most benchmarks first, then most papers with code.

Arithmetic Reasoning

5 benchmarks · 112 papers with code

Error Understanding

2 benchmarks · 5 papers with code

Emotion Interpretation

2 benchmarks · 3 papers with code

Human Judgment Correlation

2 benchmarks · 3 papers with code

Natural Language Visual Grounding

1 benchmark · 30 papers with code

Odd One Out

1 benchmark · 12 papers with code

Image Paragraph Captioning

1 benchmark · 5 papers with code

Analogical Similarity

1 benchmark · 4 papers with code

Human Judgment Classification

1 benchmark · 2 papers with code

Identify Odd Metapor

1 benchmark · 2 papers with code

Commonsense Reasoning for RL

1 benchmark · 1 paper with code

ARC

0 benchmarks · 155 papers with code

Systematic Generalization

0 benchmarks · 77 papers with code

X-ray Classification

0 benchmarks · 27 papers with code

Discrete Choice Models

0 benchmarks · 16 papers with code

Causal Identification

0 benchmarks · 14 papers with code

Abstract Argumentation

0 benchmarks · 6 papers with code

Lightweight Deployment

0 benchmarks · 6 papers with code

Theory of Mind Modeling

0 benchmarks · 6 papers with code

Music Genre Transfer

0 benchmarks · 4 papers with code

Assortment Optimization

0 benchmarks · 3 papers with code

Pre-election ratings estimation

0 benchmarks · 1 paper with code

When should a hot water tank be replaced?

0 benchmarks · 1 paper with code

4 tasks in Reasoning are filed only under a parent task from another area and are not listed on this page; the parent's task page carries them.

Task tree and counts are the archive's, frozen 2025-07-28 archive 2025-07-28. Nothing here is re-ranked.