{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/lanidenn-multilingual-language-identification","title":"LanideNN: Multilingual Language Identification on Character Window","arxiv_id":"1701.03338","date":"2017-01-12","proceeding":"EACL 2017 4","authors":["Tom Kocmi","Ondřej Bojar"],"abstract":"In language identification, a common first step in natural language\nprocessing, we want to automatically determine the language of some input text.\nMonolingual language identification assumes that the given document is written\nin one language. In multilingual language identification, the document is\nusually in two or three languages and we just want their names. We aim one step\nfurther and propose a method for textual language identification where\nlanguages can change arbitrarily and the goal is to identify the spans of each\nof the languages. Our method is based on Bidirectional Recurrent Neural\nNetworks and it performs well in monolingual and multilingual language\nidentification tasks on six datasets covering 131 languages. The method keeps\nthe accuracy also for short documents and across domains, so it is ideal for\noff-the-shelf use without preparation of training data.","url_abs":"http://arxiv.org/abs/1701.03338v2","url_pdf":"http://arxiv.org/pdf/1701.03338v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"lanidenn-multilingual-language-identification","repo_url":"https://github.com/tomkocmi/LanideNN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"language-identification","task_name":"Language Identification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1701.03338","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}