{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/language-modeling-with-sparse-product-of","title":"Language Modeling with Sparse Product of Sememe Experts","arxiv_id":"1810.12387","date":"2018-10-29","proceeding":"EMNLP 2018 10","authors":["Yihong Gu","Jun Yan","Hao Zhu","Zhiyuan Liu","Ruobing Xie","Maosong Sun","Fen Lin","Leyu Lin"],"abstract":"Most language modeling methods rely on large-scale data to statistically\nlearn the sequential patterns of words. In this paper, we argue that words are\natomic language units but not necessarily atomic semantic units. Inspired by\nHowNet, we use sememes, the minimum semantic units in human languages, to\nrepresent the implicit semantics behind words for language modeling, named\nSememe-Driven Language Model (SDLM). More specifically, to predict the next\nword, SDLM first estimates the sememe distribution gave textual context.\nAfterward, it regards each sememe as a distinct semantic expert, and these\nexperts jointly identify the most probable senses and the corresponding word.\nIn this way, SDLM enables language models to work beyond word-level\nmanipulation to fine-grained sememe-level semantics and offers us more powerful\ntools to fine-tune language models and improve the interpretability as well as\nthe robustness of language models. Experiments on language modeling and the\ndownstream application of headline gener- ation demonstrate the significant\neffect of SDLM. Source code and data used in the experiments can be accessed at\nhttps:// github.com/thunlp/SDLM-pytorch.","url_abs":"http://arxiv.org/abs/1810.12387v1","url_pdf":"http://arxiv.org/pdf/1810.12387v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"language-modeling-with-sparse-product-of","repo_url":"https://github.com/thunlp/SDLM-pytorch","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[{"method_slug":"interpretability","method_name":"Interpretability"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1810.12387","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}