{"url":"/dataset/riddle-sense","name":"RiddleSense","full_name":null,"description_markdown":"Question: I have five fingers but I am not alive. What am I? Answer: a glove.\r\n\r\nAnswering such a riddle-style question is a challenging cognitive process, in that it requires complex commonsense reasoning abilities, an understanding of figurative language, and counterfactual reasoning skills, which are all important abilities for advanced natural language understanding (NLU). However, there is currently no dedicated datasets aiming to test these abilities. Herein, we present RiddleSense, a new multiple-choice question answering task, which comes with the first large dataset (5.7k examples) for answering riddle-style commonsense questions. We systematically evaluate a wide range of models over the challenge, and point out that there is a large gap between the best-supervised model and human performance — suggesting intriguing future research in the direction of higher-order commonsense reasoning and linguistic creativity towards building advanced NLU systems.","description_withheld":null,"homepage":"https://inklab.usc.edu/RiddleSense/","introduced_date":"2021-01-02","introduced_date_note":null,"introduced_by":{"paper":"/paper/riddlesense-answering-riddle-questions-as","title":"RiddleSense: Reasoning about Riddle Questions Featuring Linguistic Creativity and Commonsense Knowledge","first_author":"Bill Yuchen Lin","url":null},"license":null,"modalities":[],"tasks":[{"name":"Common Sense Reasoning","url":"/task/common-sense-reasoning","datasets_with_task":"/datasets/task/common-sense-reasoning"},{"name":"Riddle Sense","url":"/task/riddle-sense","datasets_with_task":"/datasets/task/riddle-sense"}],"languages":[],"variants":["RiddleSense","pszemraj/riddlesense_plusplus"],"data_loaders":[],"num_papers_in_archive":23,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/riddle-sense-on-riddle-sense","task":"Riddle Sense","dataset_variant":"RiddleSense","rows":3,"metrics":["Accuracy (%)"],"first_row_in_archive_order":{"model":"DRAGON","paper":"/paper/deep-bidirectional-language-knowledge-graph","metrics":{"Accuracy (%)":"71.3"},"code_links":[{"title":"michiyasunaga/dragon","url":"https://github.com/michiyasunaga/dragon"},{"title":"HaochenLiu2000/QAP","url":"https://github.com/HaochenLiu2000/QAP"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/deep-bidirectional-language-knowledge-graph","title":"Deep Bidirectional Language-Knowledge Graph Pretraining","date":"2022-10-17","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":18,"samples_ran":2,"samples_unverified":16,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/qa-gnn-reasoning-with-language-models-and","title":"QA-GNN: Reasoning with Language Models and Knowledge Graphs for Question Answering","date":"2021-04-13","rows_on_this_dataset":1,"code_links":6,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":25,"samples_ran":1,"samples_unverified":24,"pointer_only_for_licence":4,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":2,"samples_harvested":43,"samples_ran":3,"samples_unverified":40,"pointer_only_for_licence":4,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}