{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/urdu-word-segmentation-using-conditional","title":"Urdu Word Segmentation using Conditional Random Fields (CRFs)","arxiv_id":"1806.05432","date":"2018-06-14","proceeding":"COLING 2018 8","authors":["Haris Bin Zia","Agha Ali Raza","Awais Athar"],"abstract":"State-of-the-art Natural Language Processing algorithms rely heavily on\nefficient word segmentation. Urdu is amongst languages for which word\nsegmentation is a complex task as it exhibits space omission as well as space\ninsertion issues. This is partly due to the Arabic script which although\ncursive in nature, consists of characters that have inherent joining and\nnon-joining attributes regardless of word boundary. This paper presents a word\nsegmentation system for Urdu which uses a Conditional Random Field sequence\nmodeler with orthographic, linguistic and morphological features. Our proposed\nmodel automatically learns to predict white space as word boundary as well as\nZero Width Non-Joiner (ZWNJ) as sub-word boundary. Using a manually annotated\ncorpus, our model achieves F1 score of 0.97 for word boundary identification\nand 0.85 for sub-word boundary identification tasks. We have made our code and\ncorpus publicly available to make our results reproducible.","url_abs":"http://arxiv.org/abs/1806.05432v1","url_pdf":"http://arxiv.org/pdf/1806.05432v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"urdu-word-segmentation-using-conditional","repo_url":"https://github.com/harisbinzia/Urdu-Word-Segmentation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"segmentation","task_name":"Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}