{"url":"/dataset/commonsenseqa","name":"CommonsenseQA","full_name":"CSQA","description_markdown":"The **CommonsenseQA** is a dataset for commonsense question answering task. The dataset consists of 12,247 questions with 5 choices each.\r\nThe dataset was generated by Amazon Mechanical Turk workers in the following process (an example is provided in parentheses):\r\n\r\n1. a crowd worker observes a source concept from ConceptNet (“River”) and three target concepts (“Waterfall”, “Bridge”, “Valley”) that are all related by the same ConceptNet relation (“AtLocation”),\r\n2. the worker authors three questions, one per target concept, such that only that particular target concept is the answer, while the other two distractor concepts are not, (“Where on a river can you hold a cup upright to catch water on a sunny day?”, “Where can I stand on a river to see water falling without getting wet?”, “I’m crossing the river, my feet are wet but my body is dry, where am I?”)\r\n3. for each question, another worker chooses one additional distractor from Concept Net (“pebble”, “stream”, “bank”), and the author another distractor (“mountain”, “bottom”, “island”) manually.\r\n\r\nSource: [CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge](https://paperswithcode.com/paper/commonsenseqa-a-question-answering-challenge/)\r\nImage Source: [CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge](https://paperswithcode.com/paper/commonsenseqa-a-question-answering-challenge/)","description_withheld":null,"homepage":"https://www.tau-nlp.org/commonsenseqa","introduced_date":"2019-01-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/commonsenseqa-a-question-answering-challenge","title":"CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge","first_author":"Alon Talmor","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Common Sense Reasoning","url":"/task/common-sense-reasoning","datasets_with_task":"/datasets/task/common-sense-reasoning"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["CommonsenseQA"],"data_loaders":[{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/commonsense_qa","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/rizquuula/commonsense_qa-ID","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/tau/commonsense_qa","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/huggingface/datasets","url":"https://huggingface.co/datasets/Sadanto3933/commonsense_qa","frameworks":["tf","pytorch","jax"]},{"repo":"https://github.com/facebookresearch/ParlAI","url":"https://parl.ai/docs/tasks.html#commonsenseqa","frameworks":["pytorch"]},{"repo":"https://github.com/allenai/allennlp-models","url":"https://docs.allennlp.org/models/main/models/mc/dataset_readers/commonsenseqa/","frameworks":["pytorch"]}],"num_papers_in_archive":483,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/common-sense-reasoning-on-commonsenseqa","task":"Common Sense Reasoning","dataset_variant":"CommonsenseQA","rows":38,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"GPT-4o (HPT)","paper":"/paper/hierarchical-prompting-taxonomy-a-universal","metrics":{"Accuracy":"92.54"},"code_links":[{"title":"devichand579/HPT","url":"https://github.com/devichand579/HPT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/hierarchical-prompting-taxonomy-a-universal","title":"Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles","date":"2024-06-18","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/palm-2-technical-report-1","title":"PaLM 2 Technical Report","date":"2023-05-17","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/bloomberggpt-a-large-language-model-for","title":"BloombergGPT: A Large Language Model for Finance","date":"2023-03-30","rows_on_this_dataset":4,"code_links":2,"syntology":null},{"paper":"/paper/grapeqa-graph-augmentation-and-pruning-to","title":"GrapeQA: GRaph Augmentation and Pruning to Enhance Question-Answering","date":"2023-03-22","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/deep-bidirectional-language-knowledge-graph","title":"Deep Bidirectional Language-Knowledge Graph Pretraining","date":"2022-10-17","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":18,"samples_ran":2,"samples_unverified":16,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/unifying-language-learning-paradigms","title":"UL2: Unifying Language Learning Paradigms","date":"2022-05-10","rows_on_this_dataset":3,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":16,"samples_ran":0,"samples_unverified":16,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/star-bootstrapping-reasoning-with-reasoning","title":"STaR: Bootstrapping Reasoning With Reasoning","date":"2022-03-28","rows_on_this_dataset":6,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":0,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/chain-of-thought-prompting-elicits-reasoning","title":"Chain-of-Thought Prompting Elicits Reasoning in Large Language Models","date":"2022-01-28","rows_on_this_dataset":1,"code_links":19,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":2,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/human-parity-on-commonsenseqa-augmenting-self","title":"Human Parity on CommonsenseQA: Augmenting Self-Attention with External Attention","date":"2021-12-06","rows_on_this_dataset":3,"code_links":2,"syntology":null},{"paper":"/paper/qa-gnn-reasoning-with-language-models-and","title":"QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering","date":"2021-04-13","rows_on_this_dataset":1,"code_links":6,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":25,"samples_ran":1,"samples_unverified":24,"pointer_only_for_licence":4,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/unicorn-on-rainbow-a-universal-commonsense","title":"UNICORN on RAINBOW: A Universal Commonsense Reasoning Model on a New Multitask Benchmark","date":"2021-03-24","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":0,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/muppet-massive-multi-task-representations","title":"Muppet: Massive Multi-task Representations with Pre-Finetuning","date":"2021-01-26","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/fusing-context-into-knowledge-graph-for","title":"Fusing Context Into Knowledge Graph for Commonsense Question Answering","date":"2020-12-09","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/unifiedqa-crossing-format-boundaries-with-a","title":"UnifiedQA: Crossing Format Boundaries With a Single QA System","date":"2020-05-02","rows_on_this_dataset":5,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":4,"samples_unverified":3,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/towards-generalizable-neuro-symbolic-systems","title":"Towards Generalizable Neuro-Symbolic Systems for Commonsense Question Answering","date":"2019-10-30","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/albert-a-lite-bert-for-self-supervised","title":"ALBERT: A Lite BERT for Self-supervised Learning of Language Representations","date":"2019-09-26","rows_on_this_dataset":1,"code_links":48,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":126,"samples_ran":46,"samples_unverified":80,"pointer_only_for_licence":22,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/graph-based-reasoning-over-heterogeneous","title":"Graph-Based Reasoning over Heterogeneous External Knowledge for Commonsense Question Answering","date":"2019-09-09","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":17,"samples_ran":7,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/kagnet-knowledge-aware-graph-networks-for","title":"KagNet: Knowledge-Aware Graph Networks for Commonsense Reasoning","date":"2019-09-04","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":1,"samples_unverified":11,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/align-mask-and-select-a-simple-method-for","title":"Align, Mask and Select: A Simple Method for Incorporating Commonsense Knowledge into Language Representation Models","date":"2019-08-19","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/roberta-a-robustly-optimized-bert-pretraining","title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","date":"2019-07-26","rows_on_this_dataset":1,"code_links":67,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":48,"samples_ran":22,"samples_unverified":26,"pointer_only_for_licence":23,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/explain-yourself-leveraging-language-models","title":"Explain Yourself! Leveraging Language Models for Commonsense Reasoning","date":"2019-06-06","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":5,"samples_unverified":0,"pointer_only_for_licence":1,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/commonsenseqa-a-question-answering-challenge","title":"CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge","date":"2018-11-02","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":0,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":13,"samples_harvested":291,"samples_ran":90,"samples_unverified":201,"pointer_only_for_licence":51,"papers_with_no_sample_that_ran":4,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}