{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/who-did-what-a-large-scale-person-centered","title":"Who did What: A Large-Scale Person-Centered Cloze Dataset","arxiv_id":"1608.05457","date":"2016-08-19","proceeding":"EMNLP 2016 11","authors":["Takeshi Onishi","Hai Wang","Mohit Bansal","Kevin Gimpel","David Mcallester"],"abstract":"We have constructed a new \"Who-did-What\" dataset of over 200,000\nfill-in-the-gap (cloze) multiple choice reading comprehension problems\nconstructed from the LDC English Gigaword newswire corpus. The WDW dataset has\na variety of novel features. First, in contrast with the CNN and Daily Mail\ndatasets (Hermann et al., 2015) we avoid using article summaries for question\nformation. Instead, each problem is formed from two independent articles --- an\narticle given as the passage to be read and a separate article on the same\nevents used to form the question. Second, we avoid anonymization --- each\nchoice is a person named entity. Third, the problems have been filtered to\nremove a fraction that are easily solved by simple baselines, while remaining\n84% solvable by humans. We report performance benchmarks of standard systems\nand propose the WDW dataset as a challenge task for the community.","url_abs":"http://arxiv.org/abs/1608.05457v1","url_pdf":"http://arxiv.org/pdf/1608.05457v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"articles","task_name":"Articles"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"reading-comprehension","task_name":"Reading Comprehension"}],"methods":[],"datasets_introduced":[{"slug":"who-did-what","name":"Who-did-What","full_name":"Who did What"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1608.05457","atlas_url":"https://app.syntology.ai/?focus=1608.05457","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}