{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/repartitioning-of-the-complexwebquestions","title":"Repartitioning of the ComplexWebQuestions Dataset","arxiv_id":"1807.09623","date":"2018-07-25","proceeding":null,"authors":["Alon Talmor","Jonathan Berant"],"abstract":"Recently, Talmor and Berant (2018) introduced ComplexWebQuestions - a dataset\nfocused on answering complex questions by decomposing them into a sequence of\nsimpler questions and extracting the answer from retrieved web snippets. In\ntheir work the authors used a pre-trained reading comprehension (RC) model\n(Salant and Berant, 2018) to extract the answer from the web snippets. In this\nshort note we show that training a RC model directly on the training data of\nComplexWebQuestions reveals a leakage from the training set to the test set\nthat allows to obtain unreasonably high performance. As a solution, we\nconstruct a new partitioning of ComplexWebQuestions that does not suffer from\nthis leakage and publicly release it. We also perform an empirical evaluation\non these two datasets and show that training a RC model on the training data\nsubstantially improves state-of-the-art performance.","url_abs":"http://arxiv.org/abs/1807.09623v1","url_pdf":"http://arxiv.org/pdf/1807.09623v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"repartitioning-of-the-complexwebquestions","repo_url":"https://github.com/happen2me/complex-web-questions-dataset","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"reading-comprehension","task_name":"Reading Comprehension"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1807.09623","atlas_url":"https://app.syntology.ai/?focus=1807.09623","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}