{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improved-text-language-identification-for-the","title":"Improved Text Language Identification for the South African Languages","arxiv_id":"1711.00247","date":"2017-11-01","proceeding":null,"authors":["Bernardt Duvenhage","Mfundo Ntini","Phala Ramonyai"],"abstract":"Virtual assistants and text chatbots have recently been gaining popularity.\nGiven the short message nature of text-based chat interactions, the language\nidentification systems of these bots might only have 15 or 20 characters to\nmake a prediction. However, accurate text language identification is important,\nespecially in the early stages of many multilingual natural language processing\npipelines.\n  This paper investigates the use of a naive Bayes classifier, to accurately\npredict the language family that a piece of text belongs to, combined with a\nlexicon based classifier to distinguish the specific South African language\nthat the text is written in. This approach leads to a 31% reduction in the\nlanguage detection error.\n  In the spirit of reproducible research the training and testing datasets as\nwell as the code are published on github. Hopefully it will be useful to create\na text language identification shared task for South African languages.","url_abs":"http://arxiv.org/abs/1711.00247v1","url_pdf":"http://arxiv.org/pdf/1711.00247v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improved-text-language-identification-for-the","repo_url":"https://github.com/praekelt/feersum-lid-shared-task","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"language-identification","task_name":"Language Identification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1711.00247","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}