{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/building-a-comprehensive-syntactic-and","title":"Building a comprehensive syntactic and semantic corpus of Chinese clinical texts","arxiv_id":"1611.02091","date":"2016-11-07","proceeding":null,"authors":["Bin He","Bin Dong","Yi Guan","Jinfeng Yang","Zhipeng Jiang","Qiubin Yu","Jianyi Cheng","Chunyan Qu"],"abstract":"Objective: To build a comprehensive corpus covering syntactic and semantic\nannotations of Chinese clinical texts with corresponding annotation guidelines\nand methods as well as to develop tools trained on the annotated corpus, which\nsupplies baselines for research on Chinese texts in the clinical domain.\n  Materials and methods: An iterative annotation method was proposed to train\nannotators and to develop annotation guidelines. Then, by using annotation\nquality assurance measures, a comprehensive corpus was built, containing\nannotations of part-of-speech (POS) tags, syntactic tags, entities, assertions,\nand relations. Inter-annotator agreement (IAA) was calculated to evaluate the\nannotation quality and a Chinese clinical text processing and information\nextraction system (CCTPIES) was developed based on our annotated corpus.\n  Results: The syntactic corpus consists of 138 Chinese clinical documents with\n47,424 tokens and 2553 full parsing trees, while the semantic corpus includes\n992 documents that annotated 39,511 entities with their assertions and 7695\nrelations. IAA evaluation shows that this comprehensive corpus is of good\nquality, and the system modules are effective.\n  Discussion: The annotated corpus makes a considerable contribution to natural\nlanguage processing (NLP) research into Chinese texts in the clinical domain.\nHowever, this corpus has a number of limitations. Some additional types of\nclinical text should be introduced to improve corpus coverage and active\nlearning methods should be utilized to promote annotation efficiency.\n  Conclusions: In this study, several annotation guidelines and an annotation\nmethod for Chinese clinical texts were proposed, and a comprehensive corpus\nwith its NLP modules were constructed, providing a foundation for further study\nof applying NLP techniques to Chinese texts in the clinical domain.","url_abs":"http://arxiv.org/abs/1611.02091v2","url_pdf":"http://arxiv.org/pdf/1611.02091v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"building-a-comprehensive-syntactic-and","repo_url":"https://github.com/WILAB-HIT/Resources","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"active-learning","task_name":"Active Learning"},{"task_slug":"pos","task_name":"POS"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}