{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/swag-a-large-scale-adversarial-dataset-for","title":"SWAG: A Large-Scale Adversarial Dataset for Grounded Commonsense Inference","arxiv_id":"1808.05326","date":"2018-08-16","proceeding":"EMNLP 2018 10","authors":["Rowan Zellers","Yonatan Bisk","Roy Schwartz","Yejin Choi"],"abstract":"Given a partial description like \"she opened the hood of the car,\" humans can\nreason about the situation and anticipate what might come next (\"then, she\nexamined the engine\"). In this paper, we introduce the task of grounded\ncommonsense inference, unifying natural language inference and commonsense\nreasoning.\n  We present SWAG, a new dataset with 113k multiple choice questions about a\nrich spectrum of grounded situations. To address the recurring challenges of\nthe annotation artifacts and human biases found in many existing datasets, we\npropose Adversarial Filtering (AF), a novel procedure that constructs a\nde-biased dataset by iteratively training an ensemble of stylistic classifiers,\nand using them to filter the data. To account for the aggressive adversarial\nfiltering, we use state-of-the-art language models to massively oversample a\ndiverse set of potential counterfactuals. Empirical results demonstrate that\nwhile humans can solve the resulting inference problems with high accuracy\n(88%), various competitive models struggle on our task. We provide\ncomprehensive analysis that indicates significant opportunities for future\nresearch.","url_abs":"http://arxiv.org/abs/1808.05326v1","url_pdf":"http://arxiv.org/pdf/1808.05326v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"multiple-choice","task_name":"Multiple-choice"},{"task_slug":"natural-language-inference","task_name":"Natural Language Inference"},{"task_slug":"question-answering","task_name":"Question Answering"}],"methods":[],"datasets_introduced":[{"slug":"swag","name":"SWAG","full_name":"Situations With Adversarial Generations"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/common-sense-reasoning-on-swag","task":"Common Sense Reasoning","dataset":"SWAG","model":"ESIM + ELMo","rank_in_archive_order":4,"of":5,"metrics":{"Dev":"59.1","Test":"59.2"},"uses_additional_data":false},{"leaderboard":"/sota/common-sense-reasoning-on-swag","task":"Common Sense Reasoning","dataset":"SWAG","model":"ESIM + GloVe","rank_in_archive_order":5,"of":5,"metrics":{"Dev":"51.9","Test":"52.7"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1808.05326","atlas_url":"https://app.syntology.ai/?focus=1808.05326","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}