{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-reproduction-of-apple-s-bi-directional-lstm","title":"A reproduction of Apple's bi-directional LSTM models for language identification in short strings","arxiv_id":"2102.06282","date":"2021-02-11","proceeding":"EACL 2021 2","authors":["Mads Toftrup","Søren Asger Sørensen","Manuel R. Ciosici","Ira Assent"],"abstract":"Language Identification is the task of identifying a document's language. For applications like automatic spell checker selection, language identification must use very short strings such as text message fragments. In this work, we reproduce a language identification architecture that Apple briefly sketched in a blog post. We confirm the bi-LSTM model's performance and find that it outperforms current open-source language identifiers. We further find that its language identification mistakes are due to confusion between related languages.","url_abs":"https://arxiv.org/abs/2102.06282v1","url_pdf":"https://arxiv.org/pdf/2102.06282v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-reproduction-of-apple-s-bi-directional-lstm","repo_url":"https://github.com/AU-DIS/LSTM_langid","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"language-identification","task_name":"Language Identification"}],"methods":[{"method_slug":"bilstm","method_name":"BiLSTM"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/language-identification-on-opensubtitles","task":"Language Identification","dataset":"OpenSubtitles","model":"Apple bi-LSTM","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"91.37"},"uses_additional_data":false},{"leaderboard":"/sota/language-identification-on-universal","task":"Language Identification","dataset":"Universal Dependencies","model":"Apple bi-LSTM","rank_in_archive_order":1,"of":1,"metrics":{"Accuracy":"86.93"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2102.06282","atlas_url":"https://app.syntology.ai/?focus=2102.06282","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}