{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/xtreme-r-towards-more-challenging-and-nuanced","title":"XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation","arxiv_id":"2104.07412","date":"2021-04-15","proceeding":"EMNLP 2021 11","authors":["Sebastian Ruder","Noah Constant","Jan Botha","Aditya Siddhant","Orhan Firat","Jinlan Fu","PengFei Liu","Junjie Hu","Dan Garrette","Graham Neubig","Melvin Johnson"],"abstract":"Machine learning has brought striking advances in multilingual natural language processing capabilities over the past year. For example, the latest techniques have improved the state-of-the-art performance on the XTREME multilingual benchmark by more than 13 points. While a sizeable gap to human-level performance remains, improvements have been easier to achieve in some tasks than in others. This paper analyzes the current state of cross-lingual transfer learning and summarizes some lessons learned. In order to catalyze meaningful progress, we extend XTREME to XTREME-R, which consists of an improved set of ten natural language understanding tasks, including challenging language-agnostic retrieval tasks, and covers 50 typologically diverse languages. In addition, we provide a massively multilingual diagnostic suite (MultiCheckList) and fine-grained multi-dataset evaluation capabilities through an interactive public leaderboard to gain a better understanding of such models. The leaderboard and code for XTREME-R will be made available at https://sites.research.google/xtreme and https://github.com/google-research/xtreme respectively.","url_abs":"https://arxiv.org/abs/2104.07412v2","url_pdf":"https://arxiv.org/pdf/2104.07412v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"xtreme-r-towards-more-challenging-and-nuanced","repo_url":"https://github.com/google-research/xtreme","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"cross-lingual-transfer","task_name":"Cross-Lingual Transfer"},{"task_slug":"diagnostic","task_name":"Diagnostic"},{"task_slug":"natural-language-understanding","task_name":"Natural Language Understanding"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2104.07412","atlas_url":"https://app.syntology.ai/?focus=2104.07412","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}