{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sensebert-driving-some-sense-into-bert","title":"SenseBERT: Driving Some Sense into BERT","arxiv_id":"1908.05646","date":"2019-08-15","proceeding":"ACL 2020 6","authors":["Yoav Levine","Barak Lenz","Or Dagan","Ori Ram","Dan Padnos","Or Sharir","Shai Shalev-Shwartz","Amnon Shashua","Yoav Shoham"],"abstract":"The ability to learn from large unlabeled corpora has allowed neural language models to advance the frontier in natural language understanding. However, existing self-supervision techniques operate at the word form level, which serves as a surrogate for the underlying semantic content. This paper proposes a method to employ weak-supervision directly at the word sense level. Our model, named SenseBERT, is pre-trained to predict not only the masked words but also their WordNet supersenses. Accordingly, we attain a lexical-semantic level language model, without the use of human annotation. SenseBERT achieves significantly improved lexical understanding, as we demonstrate by experimenting on SemEval Word Sense Disambiguation, and by attaining a state of the art result on the Word in Context task.","url_abs":"https://arxiv.org/abs/1908.05646v2","url_pdf":"https://arxiv.org/pdf/1908.05646v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"},{"task_slug":"natural-language-inference","task_name":"Natural Language Inference"},{"task_slug":"natural-language-understanding","task_name":"Natural Language Understanding"},{"task_slug":"word-sense-disambiguation","task_name":"Word Sense Disambiguation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/natural-language-inference-on-qnli","task":"Natural Language Inference","dataset":"QNLI","model":"SenseBERT-base 110M","rank_in_archive_order":33,"of":43,"metrics":{"Accuracy":"90.6%"},"uses_additional_data":false},{"leaderboard":"/sota/natural-language-inference-on-rte","task":"Natural Language Inference","dataset":"RTE","model":"SenseBERT-base 110M","rank_in_archive_order":62,"of":90,"metrics":{"Accuracy":"67.5%"},"uses_additional_data":false},{"leaderboard":"/sota/word-sense-disambiguation-on-words-in-context","task":"Word Sense Disambiguation","dataset":"Words in Context","model":"SenseBERT-large 340M","rank_in_archive_order":11,"of":37,"metrics":{"Accuracy":"72.1"},"uses_additional_data":false},{"leaderboard":"/sota/word-sense-disambiguation-on-words-in-context","task":"Word Sense Disambiguation","dataset":"Words in Context","model":"SenseBERT-base 110M","rank_in_archive_order":12,"of":37,"metrics":{"Accuracy":"70.3"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1908.05646","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}