{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/are-you-tough-enough-framework-for-robustness","title":"Are you tough enough? Framework for Robustness Validation of Machine Comprehension Systems","arxiv_id":"1812.02205","date":"2018-12-05","proceeding":null,"authors":["Barbara Rychalska","Dominika Basaj","Przemyslaw Biecek"],"abstract":"Deep Learning NLP domain lacks procedures for the analysis of model\nrobustness. In this paper we propose a framework which validates robustness of\nany Question Answering model through model explainers. We propose that a robust\nmodel should transgress the initial notion of semantic similarity induced by\nword embeddings to learn a more human-like understanding of meaning. We test\nthis property by manipulating questions in two ways: swapping important\nquestion word for 1) its semantically correct synonym and 2) for word vector\nthat is close in embedding space. We estimate importance of words in asked\nquestions with Locally Interpretable Model Agnostic Explanations method (LIME).\nWith these two steps we compare state-of-the-art Q&A models. We show that\nalthough accuracy of state-of-the-art models is high, they are very fragile to\nchanges in the input. Moreover, we propose 2 adversarial training scenarios\nwhich raise model sensitivity to true synonyms by up to 7% accuracy measure.\nOur findings help to understand which models are more stable and how they can\nbe improved. In addition, we have created and published a new dataset that may\nbe used for validation of robustness of a Q&A model.","url_abs":"http://arxiv.org/abs/1812.02205v1","url_pdf":"http://arxiv.org/pdf/1812.02205v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"are-you-tough-enough-framework-for-robustness","repo_url":"https://github.com/MI2DataLab/nlp_interpretability_framework","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"reading-comprehension","task_name":"Reading Comprehension"},{"task_slug":"semantic-similarity","task_name":"Semantic Similarity"},{"task_slug":"semantic-textual-similarity","task_name":"Semantic Textual Similarity"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}