{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-chinese-dataset-with-negative-full-forms","title":"A Chinese Dataset with Negative Full Forms for General Abbreviation Prediction","arxiv_id":"1712.06289","date":"2017-12-18","proceeding":"LREC 2018 5","authors":["Yi Zhang","Xu sun"],"abstract":"Abbreviation is a common phenomenon across languages, especially in Chinese.\nIn most cases, if an expression can be abbreviated, its abbreviation is used\nmore often than its fully expanded forms, since people tend to convey\ninformation in a most concise way. For various language processing tasks,\nabbreviation is an obstacle to improving the performance, as the textual form\nof an abbreviation does not express useful information, unless it's expanded to\nthe full form. Abbreviation prediction means associating the fully expanded\nforms with their abbreviations. However, due to the deficiency in the\nabbreviation corpora, such a task is limited in current studies, especially\nconsidering general abbreviation prediction should also include those full form\nexpressions that do not have valid abbreviations, namely the negative full\nforms (NFFs). Corpora incorporating negative full forms for general\nabbreviation prediction are few in number. In order to promote the research in\nthis area, we build a dataset for general Chinese abbreviation prediction,\nwhich needs a few preprocessing steps, and evaluate several different models on\nthe built dataset. The dataset is available at\nhttps://github.com/lancopku/Chinese-abbreviation-dataset","url_abs":"http://arxiv.org/abs/1712.06289v1","url_pdf":"http://arxiv.org/pdf/1712.06289v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-chinese-dataset-with-negative-full-forms","repo_url":"https://github.com/lancopku/Chinese-abbreviation-dataset","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"form","task_name":"Form"},{"task_slug":"prediction","task_name":"Prediction"},{"task_slug":null,"task_name":"valid"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}