{"url":"/method/xlsr","slug":"xlsr","name":"XLSR","full_name":"XLSR","full_name_withheld":false,"description_markdown":"**XLSR** is a multilingual speech recognition model built on wav2vec 2.0 which is trained by solving a contrastive task over masked latent speech representations and jointly learns a quantization of the latents shared across languages. The model is fine-tuned on labeled data and experiments show that cross-lingual pretraining significantly outperforms monolingual pretraining. A shared quantization module over feature encoder representations produces multilingual quantized speech units whose embeddings are then used as targets for a [Transformer](https://paperswithcode.com/method/transformer) trained by contrastive learning. The model learns to share discrete tokens across languages, creating bridges across languages.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/2006.13979v2","title":"Unsupervised Cross-lingual Representation Learning for Speech Recognition","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Speech Recognition","url":"/methods/category/speech-recognition","pwc_aliases":[]}],"n_papers_tagged":20,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"EMO-Debias: Benchmarking Gender Debiasing Techniques in Multi-Label Speech Emotion Recognition","date":"2025-06-05","arxiv_id":"2506.04652","n_code_links":0,"syntology":null},{"paper":null,"title":"Self-supervised ASR Models and Features For Dysarthric and Elderly Speech Recognition","date":"2024-07-03","arxiv_id":"2407.13782","n_code_links":0,"syntology":null},{"paper":null,"title":"Less Forgetting for Better Generalization: Exploring Continual-learning Fine-tuning Methods for Speech Self-supervised Representations","date":"2024-06-30","arxiv_id":"2407.00756","n_code_links":0,"syntology":null},{"paper":null,"title":"Transcription and translation of videos using fine-tuned XLSR Wav2Vec2 on custom dataset and mBART","date":"2024-03-01","arxiv_id":"2403.00212","n_code_links":0,"syntology":null},{"paper":null,"title":"End to end Hindi to English speech conversion using Bark, mBART and a finetuned XLSR Wav2Vec2","date":"2024-01-11","arxiv_id":"2401.06183","n_code_links":0,"syntology":null},{"paper":null,"title":"Self-supervised Adaptive Pre-training of Multilingual Speech Models for Language and Dialect Identification","date":"2023-12-12","arxiv_id":"2312.07338","n_code_links":0,"syntology":null},{"paper":null,"title":"Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching","date":"2023-11-25","arxiv_id":"2311.15077","n_code_links":0,"syntology":null},{"paper":null,"title":"Modular Speech-to-Text Translation for Zero-Shot Cross-Modal Transfer","date":"2023-10-05","arxiv_id":"2310.03724","n_code_links":0,"syntology":null},{"paper":"/paper/zero-resource-code-switched-speech-benchmark","title":"Zero Resource Code-switched Speech Benchmark Using Speech Utterance Pairs For Multiple Spoken Languages","date":"2023-10-04","arxiv_id":"2310.03018","n_code_links":1,"syntology":null},{"paper":null,"title":"Indonesian Automatic Speech Recognition with XLSR-53","date":"2023-08-20","arxiv_id":"2308.11589","n_code_links":0,"syntology":null},{"paper":"/paper/allophant-cross-lingual-phoneme-recognition","title":"Allophant: Cross-lingual Phoneme Recognition with Articulatory Attributes","date":"2023-06-07","arxiv_id":"2306.04306","n_code_links":1,"syntology":null},{"paper":"/paper/improving-the-previous-state-of-the-art","title":"Improving the previous state-of-the-art Frisian ASR by fine-tuning XLS-R","date":"2023-03-31","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":"/paper/improving-spoken-language-identification-with","title":"Improving Spoken Language Identification with Map-Mix","date":"2023-02-16","arxiv_id":"2302.08229","n_code_links":1,"syntology":null},{"paper":null,"title":"Improved Self-Supervised Multilingual Speech Representation Learning Combined with Auxiliary Language Information","date":"2022-12-07","arxiv_id":"2212.03476","n_code_links":0,"syntology":null},{"paper":null,"title":"Automatic Speech Recognition of Low-Resource Languages Based on Chukchi","date":"2022-10-11","arxiv_id":"2210.05726","n_code_links":0,"syntology":null},{"paper":null,"title":"Cross-lingual Self-Supervised Speech Representations for Improved Dysarthric Speech Recognition","date":"2022-04-04","arxiv_id":"2204.01670","n_code_links":0,"syntology":null},{"paper":null,"title":"Language Adaptive Cross-lingual Speech Representation Learning with Sparse Sharing Sub-networks","date":"2022-03-09","arxiv_id":"2203.04583","n_code_links":0,"syntology":null},{"paper":null,"title":"Joint Unsupervised and Supervised Training for Multilingual ASR","date":"2021-11-15","arxiv_id":"2111.08137","n_code_links":0,"syntology":null},{"paper":null,"title":"Magic dust for cross-lingual adaptation of monolingual wav2vec-2.0","date":"2021-10-07","arxiv_id":"2110.03560","n_code_links":0,"syntology":null},{"paper":"/paper/unsupervised-cross-lingual-representation-3","title":"Unsupervised Cross-lingual Representation Learning for Speech Recognition","date":"2020-06-24","arxiv_id":"2006.13979","n_code_links":8,"syntology":{"ran":0,"of":6,"unverified":6,"pointer_only":0}}],"papers_shown":20,"tasks":[{"task":"/task/speech-recognition","name":"Speech Recognition","papers":15},{"task":"/task/speech-recognition-1","name":"speech-recognition","papers":14},{"task":"/task/automatic-speech-recognition-2","name":"Automatic Speech Recognition","papers":8},{"task":"/task/automatic-speech-recognition","name":"Automatic Speech Recognition (ASR)","papers":7},{"task":"/task/language-modeling","name":"Language Modeling","papers":4},{"task":"/task/language-modelling","name":"Language Modelling","papers":4},{"task":"/task/representation-learning","name":"Representation Learning","papers":3},{"task":"/task/translation","name":"Translation","papers":3},{"task":"/task/cross-lingual-transfer","name":"Cross-Lingual Transfer","papers":2},{"task":"/task/language-identification","name":"Language Identification","papers":2},{"task":"/task/speech-representation-learning","name":"Speech Representation Learning","papers":2},{"task":"/task/speech-to-text","name":"Speech-to-Text","papers":2},{"task":"/task/spoken-language-identification","name":"Spoken language identification","papers":2},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":2},{"task":"/task/alzheimer-s-disease-detection","name":"Alzheimer's Disease Detection","papers":1},{"task":"/task/attribute","name":"Attribute","papers":1},{"task":"/task/benchmarking","name":"Benchmarking","papers":1},{"task":"/task/continual-learning","name":"Continual Learning","papers":1},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":1},{"task":"/task/dialect-identification","name":"Dialect Identification","papers":1}],"tasks_shown":20,"n_tasks":38,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":2},{"year":"2022","papers":4},{"year":"2023","papers":8},{"year":"2024","papers":4},{"year":"2025","papers":1}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/xlsr"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}