{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/building-a-syllable-database-to-solve-the","title":"Building a Syllable Database to Solve the Problem of Khmer Word Segmentation","arxiv_id":"1703.02166","date":"2017-03-07","proceeding":null,"authors":["Nam Tran Van"],"abstract":"Word segmentation is a basic problem in natural language processing. With the\nlanguages having the complex writing system like the Khmer language in Southern\nof Vietnam, this problem really very intractable, posing the significant\nchallenges. Although there are some experts in Vietnam as well as international\nhaving deeply researched this problem, there are still no reasonable results\nmeeting the demand, in particular, no treated thoroughly the ambiguous\nphenomenon, in the process of Khmer language processing so far. This paper\npresent a solution based on the syllable division into component clusters using\ntwo syllable models proposed, thereby building a Khmer syllable database, is\nstill not actually available. This method using a lexical database updated from\nthe online Khmer dictionaries and some supported dictionaries serving role of\ntraining data and complementary linguistic characteristics. Each component\ncluster is labelled and located by the first and last letter to identify\nentirety a syllable. This approach is workable and the test results achieve\nhigh accuracy, eliminate the ambiguity, contribute to solving the problem of\nword segmentation and applying efficiency in Khmer language processing.","url_abs":"http://arxiv.org/abs/1703.02166v1","url_pdf":"http://arxiv.org/pdf/1703.02166v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"building-a-syllable-database-to-solve-the","repo_url":"https://github.com/buda-base/lucene-km","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}