Datasets › CommonsenseQA

CommonsenseQA (CSQA)

Introduced by Alon Talmor et al. in CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge1 Jan 2019 archive 2025-07-28

The CommonsenseQA is a dataset for commonsense question answering task. The dataset consists of 12,247 questions with 5 choices each. The dataset was generated by Amazon Mechanical Turk workers in the following process (an example is provided in parentheses):

  1. a crowd worker observes a source concept from ConceptNet (“River”) and three target concepts (“Waterfall”, “Bridge”, “Valley”) that are all related by the same ConceptNet relation (“AtLocation”),
  2. the worker authors three questions, one per target concept, such that only that particular target concept is the answer, while the other two distractor concepts are not, (“Where on a river can you hold a cup upright to catch water on a sunny day?”, “Where can I stand on a river to see water falling without getting wet?”, “I’m crossing the river, my feet are wet but my body is dry, where am I?”)
  3. for each question, another worker chooses one additional distractor from Concept Net (“pebble”, “stream”, “bank”), and the author another distractor (“mountain”, “bottom”, “island”) manually.

Source: CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge Image Source: CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Common Sense Reasoning CommonsenseQA GPT-4o (HPT) Accuracy 92.54 Hierarchical Prompting Taxonomy: A Universal Evaluation... devichand579/HPT 38 Compare

Papers archive 2025-07-28

22 shown of 22 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 483. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles 1 1 18 Jun 2024 not harvested
PaLM 2 Technical Report 1 1 17 May 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 4 30 Mar 2023 not harvested
GrapeQA: GRaph Augmentation and Pruning to Enhance Question-Answering 0 1 22 Mar 2023 not harvested
Deep Bidirectional Language-Knowledge Graph Pretraining 2 1 17 Oct 2022 ran 2 of 18 samples (16 unverified)
UL2: Unifying Language Learning Paradigms 2 3 10 May 2022 ran 0 of 16 samples (16 unverified)
STaR: Bootstrapping Reasoning With Reasoning 1 6 28 Mar 2022 ran 0 of 2 samples (2 unverified)
Chain-of-Thought Prompting Elicits Reasoning in Large Language Models 19 1 28 Jan 2022 ran 2 of 7 samples (5 unverified)
Human Parity on CommonsenseQA: Augmenting Self-Attention with External Attention 2 3 6 Dec 2021 not harvested
QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering 6 1 13 Apr 2021 ran 1 of 25 samples (24 unverified; 4 pointer-only for licence)
UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark 1 1 24 Mar 2021 ran 0 of 5 samples (5 unverified)
Muppet: Massive Multi-task Representations with Pre-Finetuning 2 1 26 Jan 2021 not harvested
Fusing Context Into Knowledge Graph for Commonsense Question Answering 2 1 9 Dec 2020 not harvested
UnifiedQA: Crossing Format Boundaries With a Single QA System 2 5 2 May 2020 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
Towards Generalizable Neuro-Symbolic Systems for Commonsense Question Answering 0 1 30 Oct 2019 not harvested
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations 48 1 26 Sep 2019 ran 46 of 126 samples (80 unverified; 22 pointer-only for licence)
Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering 1 1 9 Sep 2019 ran 7 of 17 samples (10 unverified)
KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning 2 1 4 Sep 2019 ran 1 of 12 samples (11 unverified)
Align, Mask and Select: A Simple Method for Incorporating Commonsense Knowledge into Language Representation Models 0 1 19 Aug 2019 not harvested
RoBERTa: A Robustly Optimized BERT Pretraining Approach 67 1 26 Jul 2019 ran 22 of 48 samples (26 unverified; 23 pointer-only for licence)
Explain Yourself! Leveraging Language Models for Commonsense Reasoning 1 1 6 Jun 2019 ran 5 of 5 samples (0 unverified; 1 pointer-only for licence)
CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge 4 1 2 Nov 2018 ran 0 of 3 samples (3 unverified)

Dataset loaders archive 2025-07-28

6 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • CommonsenseQA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections