{"url":"/dataset/winowhy","name":"WinoWhy","full_name":null,"description_markdown":"The **WinoWhy dataset** is a resource that provides human-annotated reasons for answering Winograd Schema Challenge (WSC) questions. It includes the original WSC dataset and 4095 WinoWhy reasons (15 for each WSC question) that could justify the pronoun coreference choices in WSC. \r\n\r\nThe reasons in WinoWhy come from three sources:\r\n1. Human: Reasons provided by human beings.\r\n2. Human Reverse: Human reasons for the paired WSC question.\r\n3. Generation Model: The reasons generated by GPT-2 with the same question.\r\n\r\nEach WSC question has 5 reasons from each source. These reasons are then used to categorize what types of commonsense knowledge are needed to solve the WSC question. The dataset also includes a new task called WinoWhy, which requires models to distinguish plausible reasons from very similar but wrong reasons for all WSC questions. This helps to investigate whether current WSC models can understand the commonsense or simply solve the WSC questions based on the statistical bias of the dataset.","description_withheld":null,"homepage":"https://github.com/colinzhaoust/WinoWhy","introduced_date":"2020-05-12","introduced_date_note":null,"introduced_by":{"paper":"/paper/winowhy-a-deep-diagnosis-of-essential","title":"WinoWhy: A Deep Diagnosis of Essential Commonsense Knowledge for Answering Winograd Schema Challenge","first_author":"Hongming Zhang","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["WinoWhy"],"data_loaders":[],"num_papers_in_archive":10,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}