{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/empirical-evaluation-of-character-based-model","title":"Empirical Evaluation of Character-Based Model on Neural Named-Entity Recognition in Indonesian Conversational Texts","arxiv_id":"1805.12291","date":"2018-05-31","proceeding":"WS 2018 11","authors":["Kemal Kurniawan","Samuel Louvan"],"abstract":"Despite the long history of named-entity recognition (NER) task in the\nnatural language processing community, previous work rarely studied the task on\nconversational texts. Such texts are challenging because they contain a lot of\nword variations which increase the number of out-of-vocabulary (OOV) words. The\nhigh number of OOV words poses a difficulty for word-based neural models.\nMeanwhile, there is plenty of evidence to the effectiveness of character-based\nneural models in mitigating this OOV problem. We report an empirical evaluation\nof neural sequence labeling models with character embedding to tackle NER task\nin Indonesian conversational texts. Our experiments show that (1) character\nmodels outperform word embedding-only models by up to 4 $F_1$ points, (2)\ncharacter models perform better in OOV cases with an improvement of as high as\n15 $F_1$ points, and (3) character models are robust against a very high OOV\nrate.","url_abs":"http://arxiv.org/abs/1805.12291v3","url_pdf":"http://arxiv.org/pdf/1805.12291v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"empirical-evaluation-of-character-based-model","repo_url":"https://github.com/akurniawan/pytorch-sequence-tagger","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"cg","task_name":"NER"},{"task_slug":"named-entity-recognition-1","task_name":"Named Entity Recognition"},{"task_slug":"named-entity-recognition-ner","task_name":"Named Entity Recognition (NER)"},{"task_slug":"named-entity-recognition","task_name":"named-entity-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}