{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/investigating-the-working-of-text-classifiers","title":"Investigating the Working of Text Classifiers","arxiv_id":"1801.06261","date":"2018-01-19","proceeding":"COLING 2018 8","authors":["Devendra Singh Sachan","Manzil Zaheer","Ruslan Salakhutdinov"],"abstract":"Text classification is one of the most widely studied tasks in natural\nlanguage processing. Motivated by the principle of compositionality, large\nmultilayer neural network models have been employed for this task in an attempt\nto effectively utilize the constituent expressions. Almost all of the reported\nwork train large networks using discriminative approaches, which come with a\ncaveat of no proper capacity control, as they tend to latch on to any signal\nthat may not generalize. Using various recent state-of-the-art approaches for\ntext classification, we explore whether these models actually learn to compose\nthe meaning of the sentences or still just focus on some keywords or lexicons\nfor classifying the document. To test our hypothesis, we carefully construct\ndatasets where the training and test splits have no direct overlap of such\nlexicons, but overall language structure would be similar. We study various\ntext classifiers and observe that there is a big performance drop on these\ndatasets. Finally, we show that even simple models with our proposed\nregularization techniques, which disincentivize focusing on key lexicons, can\nsubstantially improve classification accuracy.","url_abs":"http://arxiv.org/abs/1801.06261v2","url_pdf":"http://arxiv.org/pdf/1801.06261v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"investigating-the-working-of-text-classifiers","repo_url":"https://github.com/DevSinghSachan/investigating-text-classifiers","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"text-classification","task_name":"Text Classification"},{"task_slug":"text-classification-1","task_name":"text-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}