{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/glyph-aware-embedding-of-chinese-characters","title":"Glyph-aware Embedding of Chinese Characters","arxiv_id":"1709.00028","date":"2017-08-31","proceeding":"WS 2017 9","authors":["Falcon Z. Dai","Zheng Cai"],"abstract":"Given the advantage and recent success of English character-level and\nsubword-unit models in several NLP tasks, we consider the equivalent modeling\nproblem for Chinese. Chinese script is logographic and many Chinese logograms\nare composed of common substructures that provide semantic, phonetic and\nsyntactic hints. In this work, we propose to explicitly incorporate the visual\nappearance of a character's glyph in its representation, resulting in a novel\nglyph-aware embedding of Chinese characters. Being inspired by the success of\nconvolutional neural networks in computer vision, we use them to incorporate\nthe spatio-structural patterns of Chinese glyphs as rendered in raw pixels. In\nthe context of two basic Chinese NLP tasks of language modeling and word\nsegmentation, the model learns to represent each character's task-relevant\nsemantic and syntactic information in the character-level embedding.","url_abs":"http://arxiv.org/abs/1709.00028v1","url_pdf":"http://arxiv.org/pdf/1709.00028v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"glyph-aware-embedding-of-chinese-characters","repo_url":"https://github.com/falcondai/chinese-char-lm","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"language-modeling","task_name":"Language Modeling"},{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1709.00028","atlas_url":"https://app.syntology.ai/?focus=1709.00028","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}