{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/glotlid-language-identification-for-low","title":"GlotLID: Language Identification for Low-Resource Languages","arxiv_id":"2310.16248","date":"2023-10-24","proceeding":null,"authors":["Amir Hossein Kargaran","Ayyoob Imani","François Yvon","Hinrich Schütze"],"abstract":"Several recent papers have published good solutions for language identification (LID) for about 300 high-resource and medium-resource languages. However, there is no LID available that (i) covers a wide range of low-resource languages, (ii) is rigorously evaluated and reliable and (iii) efficient and easy to use. Here, we publish GlotLID-M, an LID model that satisfies the desiderata of wide coverage, reliability and efficiency. It identifies 1665 languages, a large increase in coverage compared to prior work. In our experiments, GlotLID-M outperforms four baselines (CLD3, FT176, OpenLID and NLLB) when balancing F1 and false positive rate (FPR). We analyze the unique challenges that low-resource LID poses: incorrect corpus metadata, leakage from high-resource languages, difficulty separating closely related languages, handling of macrolanguage vs varieties and in general noisy data. We hope that integrating GlotLID-M into dataset creation pipelines will improve quality and enhance accessibility of NLP technology for low-resource languages and cultures. GlotLID-M model (including future versions), code, and list of data sources are available: https://github.com/cisnlp/GlotLID.","url_abs":"https://arxiv.org/abs/2310.16248v3","url_pdf":"https://arxiv.org/pdf/2310.16248v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"glotlid-language-identification-for-low","repo_url":"https://github.com/cisnlp/glotlid","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"glotlid-language-identification-for-low","repo_url":"https://github.com/cisnlp/glotsparse","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"gone","observed_at":"2026-09-17","how":"tree_404+repo_404"}},{"paper_slug":"glotlid-language-identification-for-low","repo_url":"https://github.com/cisnlp/glotstorybook","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":{"status":"ok","spdx":"CC0-1.0"}}],"tasks":[{"task_slug":"dialect-identification","task_name":"Dialect Identification"},{"task_slug":"language-identification","task_name":"Language Identification"}],"methods":[],"datasets_introduced":[{"slug":"glotsparse","name":"GlotSparse","full_name":""},{"slug":"glotstorybook","name":"GlotStoryBook","full_name":""},{"slug":"udhr-lid","name":"udhr-lid","full_name":""}],"methods_introduced":[],"results":[{"leaderboard":"/sota/language-identification-on-glotlid-c","task":"Language Identification","dataset":"GlotLID-C","model":"GlotLID","rank_in_archive_order":1,"of":1,"metrics":{"Macro F1":"0.977"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2310.16248","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}