{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/incorporating-chinese-characters-of-words-for","title":"Incorporating Chinese Characters of Words for Lexical Sememe Prediction","arxiv_id":"1806.06349","date":"2018-06-17","proceeding":"ACL 2018 7","authors":["Huiming Jin","Hao Zhu","Zhiyuan Liu","Ruobing Xie","Maosong Sun","Fen Lin","Leyu Lin"],"abstract":"Sememes are minimum semantic units of concepts in human languages, such that\neach word sense is composed of one or multiple sememes. Words are usually\nmanually annotated with their sememes by linguists, and form linguistic\ncommon-sense knowledge bases widely used in various NLP tasks. Recently, the\nlexical sememe prediction task has been introduced. It consists of\nautomatically recommending sememes for words, which is expected to improve\nannotation efficiency and consistency. However, existing methods of lexical\nsememe prediction typically rely on the external context of words to represent\nthe meaning, which usually fails to deal with low-frequency and\nout-of-vocabulary words. To address this issue for Chinese, we propose a novel\nframework to take advantage of both internal character information and external\ncontext information of words. We experiment on HowNet, a Chinese sememe\nknowledge base, and demonstrate that our framework outperforms state-of-the-art\nbaselines by a large margin, and maintains a robust performance even for\nlow-frequency words.","url_abs":"http://arxiv.org/abs/1806.06349v1","url_pdf":"http://arxiv.org/pdf/1806.06349v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"incorporating-chinese-characters-of-words-for","repo_url":"https://github.com/thunlp/Character-enhanced-Sememe-Prediction","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"},{"task_slug":"prediction","task_name":"Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1806.06349","atlas_url":"https://app.syntology.ai/?focus=1806.06349","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}