{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/190501964","title":"Neural Chinese Named Entity Recognition via CNN-LSTM-CRF and Joint Training with Word Segmentation","arxiv_id":"1905.01964","date":"2019-04-26","proceeding":null,"authors":["Fangzhao Wu","Junxin Liu","Chuhan Wu","Yongfeng Huang","Xing Xie"],"abstract":"Chinese named entity recognition (CNER) is an important task in Chinese\nnatural language processing field. However, CNER is very challenging since\nChinese entity names are highly context-dependent. In addition, Chinese texts\nlack delimiters to separate words, making it difficult to identify the boundary\nof entities. Besides, the training data for CNER in many domains is usually\ninsufficient, and annotating enough training data for CNER is very expensive\nand time-consuming. In this paper, we propose a neural approach for CNER.\nFirst, we introduce a CNN-LSTM-CRF neural architecture to capture both local\nand long-distance contexts for CNER. Second, we propose a unified framework to\njointly train CNER and word segmentation models in order to enhance the ability\nof CNER model in identifying entity boundaries. Third, we introduce an\nautomatic method to generate pseudo labeled samples from existing labeled data\nwhich can enrich the training data. Experiments on two benchmark datasets show\nthat our approach can effectively improve the performance of Chinese named\nentity recognition, especially when training data is insufficient.","url_abs":"http://arxiv.org/abs/1905.01964v1","url_pdf":"http://arxiv.org/pdf/1905.01964v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"190501964","repo_url":"https://github.com/rxy007/cnn-lstm-crf","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"chinese-named-entity-recognition","task_name":"Chinese Named Entity Recognition"},{"task_slug":"named-entity-recognition-1","task_name":"Named Entity Recognition"},{"task_slug":"named-entity-recognition-ner","task_name":"Named Entity Recognition (NER)"},{"task_slug":"named-entity-recognition","task_name":"named-entity-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1905.01964","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}