{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/xls-r-self-supervised-cross-lingual-speech","title":"XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale","arxiv_id":"2111.09296","date":"2021-11-17","proceeding":null,"authors":["Arun Babu","Changhan Wang","Andros Tjandra","Kushal Lakhotia","Qiantong Xu","Naman Goyal","Kritika Singh","Patrick von Platen","Yatharth Saraf","Juan Pino","Alexei Baevski","Alexis Conneau","Michael Auli"],"abstract":"This paper presents XLS-R, a large-scale model for cross-lingual speech representation learning based on wav2vec 2.0. We train models with up to 2B parameters on nearly half a million hours of publicly available speech audio in 128 languages, an order of magnitude more public data than the largest known prior work. Our evaluation covers a wide range of tasks, domains, data regimes and languages, both high and low-resource. On the CoVoST-2 speech translation benchmark, we improve the previous state of the art by an average of 7.4 BLEU over 21 translation directions into English. For speech recognition, XLS-R improves over the best known prior work on BABEL, MLS, CommonVoice as well as VoxPopuli, lowering error rates by 14-34% relative on average. XLS-R also sets a new state of the art on VoxLingua107 language identification. Moreover, we show that with sufficient model size, cross-lingual pretraining can outperform English-only pretraining when translating English speech into other languages, a setting which favors monolingual pretraining. We hope XLS-R can help to improve speech processing tasks for many more languages of the world.","url_abs":"https://arxiv.org/abs/2111.09296v3","url_pdf":"https://arxiv.org/pdf/2111.09296v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"xls-r-self-supervised-cross-lingual-speech","repo_url":"https://github.com/pytorch/fairseq","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"xls-r-self-supervised-cross-lingual-speech","repo_url":"https://github.com/gatech-eic/s3-router","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"language-identification","task_name":"Language Identification"},{"task_slug":"representation-learning","task_name":"Representation Learning"},{"task_slug":"speech-recognition","task_name":"Speech Recognition"},{"task_slug":"speech-representation-learning","task_name":"Speech Representation Learning"},{"task_slug":"translation","task_name":"Translation"},{"task_slug":"speech-recognition-1","task_name":"speech-recognition"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/language-identification-on-voxlingua107-1","task":"Language Identification","dataset":"VOXLINGUA107","model":"XLS-R","rank_in_archive_order":1,"of":2,"metrics":{"Error rate":"5.7"},"uses_additional_data":false},{"leaderboard":"/sota/language-identification-on-voxlingua107-1","task":"Language Identification","dataset":"VOXLINGUA107","model":"wav2vec 2.0 LV-60K","rank_in_archive_order":2,"of":2,"metrics":{"Error rate":"7.2"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2111.09296","atlas_url":"https://app.syntology.ai/?focus=2111.09296","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}