{"url":"/method/xlm","slug":"xlm","name":"XLM","full_name":"XLM","full_name_withheld":false,"description_markdown":"**XLM** is a [Transformer](https://paperswithcode.com/method/transformer) based architecture that is pre-trained using one of three language modelling objectives:\r\n\r\n1. Causal Language Modeling - models the probability of a word given the previous words in a sentence.\r\n2. Masked Language Modeling - the masked language modeling objective of [BERT](https://paperswithcode.com/method/bert).\r\n3. Translation Language Modeling - a (new) translation language modeling objective for improving cross-lingual pre-training.\r\n\r\nThe authors find that both the CLM and MLM approaches provide strong cross-lingual features that can be used for pretraining models.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Cross-lingual Language Model Pretraining","paper":"/paper/cross-lingual-language-model-pretraining","first_author":"Guillaume Lample","n_authors":2,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/cross-lingual-language-model-pretraining"},"source":{"url":"http://arxiv.org/abs/1901.07291v1","title":"Cross-lingual Language Model Pretraining","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Autoencoding Transformers","url":"/methods/category/autoencoding-transformers","pwc_aliases":[]},{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":57,"archive_num_papers":57,"papers_newest_first":[{"paper":"/paper/biobridge-unified-bio-embedding-with-bridging","title":"BioBridge: Unified Bio-Embedding with Bridging Modality in Code-Switched EMR","date":"2024-12-16","arxiv_id":"2412.11671","n_code_links":1,"syntology":null},{"paper":null,"title":"XLM for Autonomous Driving Systems: A Comprehensive Review","date":"2024-09-16","arxiv_id":"2409.10484","n_code_links":0,"syntology":null},{"paper":"/paper/ax-to-grind-urdu-benchmark-dataset-for-urdu","title":"Ax-to-Grind Urdu: Benchmark Dataset for Urdu Fake News Detection","date":"2024-03-20","arxiv_id":"2403.14037","n_code_links":1,"syntology":null},{"paper":"/paper/mapping-transformer-leveraged-embeddings-for","title":"Mapping Transformer Leveraged Embeddings for Cross-Lingual Document Representation","date":"2024-01-12","arxiv_id":"2401.06583","n_code_links":1,"syntology":null},{"paper":null,"title":"An Empirical study of Unsupervised Neural Machine Translation: analyzing NMT output, model's behavior and sentences' contribution","date":"2023-12-19","arxiv_id":"2312.12588","n_code_links":0,"syntology":null},{"paper":null,"title":"MedAI Dialog Corpus (MEDIC): Zero-Shot Classification of Doctor and AI Responses in Health Consultations","date":"2023-10-19","arxiv_id":"2310.12489","n_code_links":0,"syntology":null},{"paper":"/paper/benchmarking-procedural-language","title":"Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish","date":"2023-09-13","arxiv_id":"2309.06698","n_code_links":1,"syntology":null},{"paper":null,"title":"PESTS: Persian_English Cross Lingual Corpus for Semantic Textual Similarity","date":"2023-05-13","arxiv_id":"2305.07893","n_code_links":0,"syntology":null},{"paper":null,"title":"CLaC at SemEval-2023 Task 2: Comparing Span-Prediction and Sequence-Labeling approaches for NER","date":"2023-05-05","arxiv_id":"2305.03845","n_code_links":0,"syntology":null},{"paper":"/paper/exploring-methods-for-building-dialects","title":"Exploring Methods for Building Dialects-Mandarin Code-Mixing Corpora: A Case Study in Taiwanese Hokkien","date":"2023-01-21","arxiv_id":"2301.08937","n_code_links":1,"syntology":null},{"paper":"/paper/align-mlm-word-embedding-alignment-is-crucial","title":"ALIGN-MLM: Word Embedding Alignment is Crucial for Multilingual Pre-training","date":"2022-11-15","arxiv_id":"2211.08547","n_code_links":1,"syntology":null},{"paper":"/paper/bert-sort-a-zero-shot-mlm-semantic-encoder-on","title":"BERT-Sort: A Zero-shot MLM Semantic Encoder on Ordinal Features for AutoML","date":"2022-06-01","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/geomlama-geo-diverse-commonsense-probing-on","title":"GeoMLAMA: Geo-Diverse Commonsense Probing on Multilingual Pre-Trained Language Models","date":"2022-05-24","arxiv_id":"2205.12247","n_code_links":1,"syntology":null},{"paper":"/paper/persian-natural-language-inference-a-meta-1","title":"Persian Natural Language Inference: A Meta-learning approach","date":"2022-05-18","arxiv_id":"2205.08755","n_code_links":1,"syntology":null},{"paper":"/paper/hiner-a-large-hindi-named-entity-recognition","title":"HiNER: A Large Hindi Named Entity Recognition Dataset","date":"2022-04-28","arxiv_id":"2204.13743","n_code_links":1,"syntology":null},{"paper":null,"title":"Team ÚFAL at CMCL 2022 Shared Task: Figuring out the correct recipe for predicting Eye-Tracking features using Pretrained Language Models","date":"2022-04-11","arxiv_id":"2204.04998","n_code_links":0,"syntology":null},{"paper":null,"title":"Are You Robert or RoBERTa? Deceiving Online Authorship Attribution Models Using Neural Text Generators","date":"2022-03-18","arxiv_id":"2203.09813","n_code_links":0,"syntology":null},{"paper":"/paper/a-passage-to-india-pre-trained-word-1","title":"\"A Passage to India\": Pre-trained Word Embeddings for Indian Languages","date":"2021-12-27","arxiv_id":"2112.13800","n_code_links":0,"syntology":null},{"paper":"/paper/prix-lm-pretraining-for-multilingual","title":"Prix-LM: Pretraining for Multilingual Knowledge Base Construction","date":"2021-10-16","arxiv_id":"2110.08443","n_code_links":1,"syntology":null},{"paper":"/paper/cross-language-learning-for-entity-matching","title":"Cross-Language Learning for Entity Matching","date":"2021-10-07","arxiv_id":"2110.03338","n_code_links":1,"syntology":null},{"paper":"/paper/turingbench-a-benchmark-environment-for","title":"TURINGBENCH: A Benchmark Environment for Turing Test in the Age of Neural Text Generation","date":"2021-09-27","arxiv_id":"2109.13296","n_code_links":3,"syntology":{"ran":0,"of":7,"unverified":7,"pointer_only":0}},{"paper":"/paper/do-images-really-do-the-talking-analysing-the","title":"Do Images really do the Talking? Analysing the significance of Images in Tamil Troll meme classification","date":"2021-08-09","arxiv_id":"2108.03886","n_code_links":1,"syntology":null},{"paper":null,"title":"PAW at SemEval-2021 Task 2: Multilingual and Cross-lingual Word-in-Context Disambiguation : Exploring Cross Lingual Transfer, Augmentations and Adversarial Training","date":"2021-08-01","arxiv_id":null,"n_code_links":0,"syntology":null},{"paper":null,"title":"A Primer on Pretrained Multilingual Language Models","date":"2021-07-01","arxiv_id":"2107.00676","n_code_links":0,"syntology":null},{"paper":null,"title":"LAWDR: Language-Agnostic Weighted Document Representations from Pre-trained Models","date":"2021-06-07","arxiv_id":"2106.03379","n_code_links":0,"syntology":null},{"paper":null,"title":"Unsupervised Multilingual Sentence Embeddings for Parallel Corpus Mining","date":"2021-05-21","arxiv_id":"2105.10419","n_code_links":0,"syntology":null},{"paper":"/paper/teamuncc-lt-edi-eacl2021-hope-speech","title":"TeamUNCC@LT-EDI-EACL2021: Hope Speech Detection using Transfer Learning with Transformers","date":"2021-04-19","arxiv_id":null,"n_code_links":1,"syntology":null},{"paper":"/paper/multilingual-language-models-predict-human","title":"Multilingual Language Models Predict Human Reading Behavior","date":"2021-04-12","arxiv_id":"2104.05433","n_code_links":1,"syntology":null},{"paper":null,"title":"Low-Resource Machine Translation Training Curriculum Fit for Low-Resource Languages","date":"2021-03-24","arxiv_id":"2103.13272","n_code_links":0,"syntology":null},{"paper":null,"title":"LightMBERT: A Simple Yet Effective Method for Multilingual BERT Distillation","date":"2021-03-11","arxiv_id":"2103.06418","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":17},{"task":"/task/translation","name":"Translation","papers":15},{"task":"/task/language-modeling","name":"Language Modeling","papers":14},{"task":"/task/sentence","name":"Sentence","papers":11},{"task":"/task/machine-translation","name":"Machine Translation","papers":10},{"task":"/task/cross-lingual-transfer","name":"Cross-Lingual Transfer","papers":8},{"task":"/task/question-answering","name":"Question Answering","papers":8},{"task":"/task/xlm-r","name":"XLM-R","papers":8},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":7},{"task":"/task/cg","name":"NER","papers":6},{"task":"/task/retrieval","name":"Retrieval","papers":6},{"task":"/task/nmt","name":"NMT","papers":5},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":5},{"task":"/task/text-classification","name":"Text Classification","papers":5},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":5},{"task":"/task/classification-1","name":"Classification","papers":4},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":4},{"task":"/task/representation-learning","name":"Representation Learning","papers":4},{"task":"/task/zero-shot-cross-lingual-transfer","name":"Zero-Shot Cross-Lingual Transfer","papers":4},{"task":"/task/text-classification-1","name":"text-classification","papers":4}],"tasks_shown":20,"n_tasks":106,"usage_by_year":[{"year":"2019","papers":9},{"year":"2020","papers":15},{"year":"2021","papers":16},{"year":"2022","papers":7},{"year":"2023","papers":6},{"year":"2024","papers":4}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/xlm"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}