{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improving-cross-domain-chinese-word","title":"Improving Cross-Domain Chinese Word Segmentation with Word Embeddings","arxiv_id":"1903.01698","date":"2019-03-05","proceeding":"NAACL 2019 6","authors":["Yuxiao Ye","Yue Zhang","Weikang Li","Likun Qiu","Jian Sun"],"abstract":"Cross-domain Chinese Word Segmentation (CWS) remains a challenge despite\nrecent progress in neural-based CWS. The limited amount of annotated data in\nthe target domain has been the key obstacle to a satisfactory performance. In\nthis paper, we propose a semi-supervised word-based approach to improving\ncross-domain CWS given a baseline segmenter. Particularly, our model only\ndeploys word embeddings trained on raw text in the target domain, discarding\ncomplex hand-crafted features and domain-specific dictionaries. Innovative\nsubsampling and negative sampling methods are proposed to derive word\nembeddings optimized for CWS. We conduct experiments on five datasets in\nspecial domains, covering domains in novels, medicine, and patent. Results show\nthat our model can obviously improve cross-domain CWS, especially in the\nsegmentation of domain-specific noun entities. The word F-measure increases by\nover 3.0% on four datasets, outperforming state-of-the-art semi-supervised and\nunsupervised cross-domain CWS approaches with a large margin. We make our code\nand data available on Github.","url_abs":"http://arxiv.org/abs/1903.01698v3","url_pdf":"http://arxiv.org/pdf/1903.01698v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improving-cross-domain-chinese-word","repo_url":"https://github.com/vatile/CWS-NAACL2019","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"chinese-word-segmentation","task_name":"Chinese Word Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"word-embeddings","task_name":"Word Embeddings"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1903.01698","atlas_url":"https://app.syntology.ai/?focus=1903.01698","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}