Home › Datasets › task › Common Sense Reasoning
Common Sense Reasoning datasets
archive 2025-07-28
56 datasets carry the task tag "Common Sense Reasoning" (the task itself: Common Sense Reasoning), ordered by the archive's paper count. Page 1 of 2: 48 shown of 56. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Common Sense Reasoning datasets 1–48 of 56
WinoGrande is a large-scale dataset of 44k problems, inspired by the original WSC design, but adjusted to improve both the scale and the hardness of the dataset.
703 papers · 7 benchmarks
The CommonsenseQA is a dataset for commonsense question answering task.
483 papers · 1 benchmark
BIG-bench (Beyond the Imitation Game Benchmark)
The Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark intended to probe large language models and extrapolate their future capabilities.
349 papers · 121 benchmarks
The AI2’s Reasoning Challenge (ARC) dataset is a multiple-choice question-answering dataset, containing questions from science exams from grade 3 to grade 9.
178 papers · 3 benchmarks
SWAG (Situations With Adversarial Generations)
Given a partial description like "she opened the hood of the car," humans can reason about the situation and anticipate what might come next ("then, she examined the engine").
163 papers · 2 benchmarks
Visual Dialog (VisDial) dataset contains human annotated questions based on images of MS COCO dataset.
159 papers · 6 benchmarks
Reading Comprehension with Commonsense Reasoning Dataset (ReCoRD) is a large-scale reading comprehension dataset which requires commonsense reasoning.
111 papers · 1 benchmark
CoS-E (Commonsense Explanations Dataset)
CoS-E consists of human explanations for commonsense reasoning in the form of natural language sequences and highlighted annotations Source: Explain Yourself!
46 papers · 0 benchmarks
IconQA (Icon Question Answering)
Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images in the daily-life context.
42 papers · 1 benchmark
PHYRE (PHYsical REasoning)
Benchmark for physical reasoning that contains a set of simple classical mechanics puzzles in a 2D physical environment.
35 papers · 2 benchmarks
A testbed for commonsense reasoning about entity knowledge, bridging fact-checking about entities with commonsense inferences.
30 papers · 0 benchmarks
CODAH (COmmonsense Dataset Adversarially-authored by Humans)
The COmmonsense Dataset Adversarially-authored by Humans (CODAH) is an evaluation set for commonsense question-answering in the sentence completion style of SWAG.
29 papers · 2 benchmarks
Useful for through two applications - automatic readability assessment and automatic text simplification.
29 papers · 0 benchmarks
XStoryCloze consists of the professionally translated version of the English StoryCloze dataset (Spring 2016 version) to 10 non-English languages.
26 papers · 0 benchmarks
MCScript is used as the official dataset of SemEval2018 Task11.
24 papers · 0 benchmarks
Moral Stories is a crowd-sourced dataset of structured narratives that describe normative and norm-divergent actions taken by individuals to accomplish certain intentions in concrete situations, and their respective consequences.
24 papers · 0 benchmarks
RecipeQA is a dataset for multimodal comprehension of cooking recipes.
24 papers · 1 benchmark
Question: I have five fingers but I am not alive.
23 papers · 1 benchmark
Event2Mind is a corpus of 25,000 event phrases covering a diverse range of everyday events and situations.
21 papers · 2 benchmarks
X-CSQA is a multilingual dataset for Commonsense reasoning research, based on CSQA.
18 papers · 0 benchmarks
A Benchmark for Robust Multi-Hop Spatial Reasoning in Texts
15 papers · 1 benchmark
Complementary Commonsense (Com2Sense) is a dataset for benchmarking commonsense reasoning ability of NLP models.
11 papers · 0 benchmarks
Housekeep a benchmark to evaluate common sense reasoning in the home for embodied AI.
11 papers · 0 benchmarks
ProtoQA is a question answering dataset for training and evaluating common sense reasoning capabilities of artificial intelligence systems in such prototypical situations.
11 papers · 0 benchmarks
TimeDial presents a crowdsourced English challenge set, for temporal commonsense reasoning, formulated as a multiple choice cloze task with around 1.5k carefully curated dialogs.
11 papers · 0 benchmarks
CC-Stories (or STORIES) is a dataset for common sense reasoning and language modeling.
10 papers · 0 benchmarks
CriticBench is a comprehensive benchmark designed to assess the abilities of Large Language Models (LLMs) to critique and rectify their reasoning across various tasks.
10 papers · 0 benchmarks
Fig-QA consists of 10256 examples of human-written creative metaphors that are paired as a Winograd schema.
10 papers · 0 benchmarks
Rainbow is multi-task benchmark for common-sense reasoning that uses different existing QA datasets: aNLI, Cosmos QA, HellaSWAG.
10 papers · 0 benchmarks
PARus (Choice of Plausible Alternatives for Russian language)
Choice of Plausible Alternatives for Russian language (PARus) evaluation provides researchers with a tool for assessing progress in open-domain commonsense causal reasoning.
7 papers · 1 benchmark
RWSD (The Winograd Schema Challenge (Russian))
A Winograd schema is a pair of sentences that differ in only one or two words and that contain an ambiguity that is resolved in opposite ways in the two sentences and requires the use of world knowledge and reasoning for its resolution.
7 papers · 1 benchmark
CITE is a crowd-sourced resource for multimodal discourse: this resource characterises inferences in image-text contexts in the domain of cooking recipes in the form of coherence relations.
6 papers · 1 benchmark
OMICS (Open Mind Indoor Common Sense)
OMICS is an extensive collection of knowledge for indoor service robots gathered from internet users.
6 papers · 0 benchmarks
SMART-101 (Simple Multimodal Algorithmic Reasoning Task Dataset)
Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, ChatGPT, etc.
6 papers · 0 benchmarks
Contains 13.6k masked-word-prediction probes, 10.5k for fine-tuning and 3.1k for testing.
5 papers · 0 benchmarks
RuCoS (Russian Reading Comprehension with Commonsense Reasoning)
Russian reading comprehension with Commonsense reasoning (RuCoS) is a large-scale reading comprehension dataset that requires commonsense reasoning.
5 papers · 1 benchmark
The Sarcasm Corpus contains sarcastic and non-sarcastic utterances of three different types, which are balanced with half of the samples being sarcastic and half non-sarcastic.
5 papers · 0 benchmarks
This dataset is collected via the WinoGAViL game to collect challenging vision-and-language associations.
5 papers · 2 benchmarks
We generate epistemic reasoning problems using modal logic to target theory of mind (tom) in natural language processing models.
4 papers · 0 benchmarks
BCOPA-CE (A Balanced COPA Test Set with cause-effect as alternatives)
We provide the BCOPA-CE test set, which has balanced token distribution in the correct and wrong alternatives and increases the difficulty of being aware of cause and effect.
3 papers · 0 benchmarks
PACS (Physical Audiovisual CommonSense) is the first audiovisual benchmark annotated for physical commonsense attributes.
3 papers · 1 benchmark
COFAR (Commonsense and Factual Reasoning in Image Search)
The COFAR (COmmonsense and FActual Reasoning) dataset is a collection of images and text queries specifically designed to challenge and evaluate image search models that aim to go beyond simple visual matching.
2 papers · 1 benchmark
CheGeKa is a Jeopardy!-like Russian QA dataset collected from the official Russian quiz database ChGK.
2 papers · 1 benchmark
G-VUE (General-purpose Visual Understanding Evaluation)
General-purpose Visual Understanding Evaluation (G-VUE) is a comprehensive benchmark covering the full spectrum of visual cognitive abilities with four functional domains -- Perceive, Ground, Reason, and Act.
2 papers · 0 benchmarks
Consists of visual arithmetic problems automatically generated using a grammar model--And-Or Graph (AOG).
2 papers · 0 benchmarks
Introduction Generalized quantifiers (e.g., few, most) are used to indicate the proportions predicates are satisfied.
2 papers · 0 benchmarks
LLMs' lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarcity of relevant data.
2 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.