{"url":"/method/xlnet","slug":"xlnet","name":"XLNet","full_name":"XLNet","full_name_withheld":false,"description_markdown":"**XLNet** is an autoregressive [Transformer](https://paperswithcode.com/method/transformer) that leverages the best of both autoregressive language modeling and autoencoding while attempting to avoid their limitations. Instead of using a fixed forward or backward factorization order as in conventional autoregressive models, XLNet maximizes the expected log likelihood of a sequence w.r.t. all possible permutations of the factorization order. Thanks to the permutation operation, the context for each position can consist of tokens from both left and right. In expectation, each position learns to utilize contextual information from all positions, i.e., capturing bidirectional context.\r\n\r\nAdditionally, inspired by the latest advancements in autogressive language modeling, XLNet integrates the segment recurrence mechanism and relative encoding scheme of [Transformer-XL](https://paperswithcode.com/method/transformer-xl) into pretraining, which empirically improves the performance especially for tasks involving a longer text sequence.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/1906.08237v2","title":"XLNet: Generalized Autoregressive Pretraining for Language Understanding","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":167,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Comparative sentiment analysis of public perception: Monkeypox vs. COVID-19 behavioral insights","date":"2025-05-12","arxiv_id":"2505.07430","n_code_links":0,"syntology":null},{"paper":null,"title":"A Character-based Diffusion Embedding Algorithm for Enhancing the Generation Quality of Generative Linguistic Steganographic Texts","date":"2025-05-02","arxiv_id":"2505.00977","n_code_links":0,"syntology":null},{"paper":null,"title":"Explainable AI for Sentiment Analysis of Human Metapneumovirus (HMPV) Using XLNet","date":"2025-02-01","arxiv_id":"2502.01663","n_code_links":0,"syntology":null},{"paper":null,"title":"Assessing Text Classification Methods for Cyberbullying Detection on Social Media Platforms","date":"2024-12-27","arxiv_id":"2412.19928","n_code_links":0,"syntology":null},{"paper":null,"title":"Feature Alignment-Based Knowledge Distillation for Efficient Compression of Large Language Models","date":"2024-12-27","arxiv_id":"2412.19449","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Multi-Class Disease Classification: Neoplasms, Cardiovascular, Nervous System, and Digestive Disorders Using Advanced LLMs","date":"2024-11-19","arxiv_id":"2411.12712","n_code_links":0,"syntology":null},{"paper":"/paper/rethinking-legal-judgement-prediction-in-a","title":"Rethinking Legal Judgement Prediction in a Realistic Scenario in the Era of Large Language Models","date":"2024-10-14","arxiv_id":"2410.10542","n_code_links":1,"syntology":null},{"paper":null,"title":"What Matters in Explanations: Towards Explainable Fake Review Detection Focusing on Transformers","date":"2024-07-24","arxiv_id":"2407.21056","n_code_links":0,"syntology":null},{"paper":"/paper/why-do-you-cite-an-investigation-on-citation","title":"Why do you cite? An investigation on citation intents and decision-making classification processes","date":"2024-07-18","arxiv_id":"2407.13329","n_code_links":0,"syntology":null},{"paper":null,"title":"Ensemble Model With Bert,Roberta and Xlnet For Molecular property prediction","date":"2024-05-30","arxiv_id":"2406.06553","n_code_links":0,"syntology":null},{"paper":null,"title":"A Hybrid Deep Learning Framework for Stock Price Prediction Considering the Investor Sentiment of Online Forum Enhanced by Popularity","date":"2024-05-17","arxiv_id":"2405.10584","n_code_links":0,"syntology":null},{"paper":null,"title":"Eliciting Personality Traits in Large Language Models","date":"2024-02-13","arxiv_id":"2402.08341","n_code_links":0,"syntology":null},{"paper":"/paper/breaking-free-transformer-models-task","title":"Breaking Free Transformer Models: Task-specific Context Attribution Promises Improved Generalizability Without Fine-tuning Pre-trained LLMs","date":"2024-01-30","arxiv_id":"2401.16638","n_code_links":1,"syntology":null},{"paper":null,"title":"An Assessment on Comprehending Mental Health through Large Language Models","date":"2024-01-09","arxiv_id":"2401.04592","n_code_links":0,"syntology":null},{"paper":null,"title":"Argumentation Element Annotation Modeling using XLNet","date":"2023-11-10","arxiv_id":"2311.06239","n_code_links":0,"syntology":null},{"paper":"/paper/modelling-sentiment-analysis-llms-and-data","title":"Modelling Sentiment Analysis: LLMs and data augmentation techniques","date":"2023-11-07","arxiv_id":"2311.04139","n_code_links":1,"syntology":null},{"paper":null,"title":"An Ensemble Method Based on the Combination of Transformers with Convolutional Neural Networks to Detect Artificially Generated Text","date":"2023-10-26","arxiv_id":"2310.17312","n_code_links":0,"syntology":null},{"paper":null,"title":"Exploring Graph Neural Networks for Indian Legal Judgment Prediction","date":"2023-10-19","arxiv_id":"2310.12800","n_code_links":0,"syntology":null},{"paper":"/paper/lexical-squad-multimodal-hate-speech-event","title":"Lexical Squad@Multimodal Hate Speech Event Detection 2023: Multimodal Hate Speech Detection using Fused Ensemble Approach","date":"2023-09-23","arxiv_id":"2309.13354","n_code_links":1,"syntology":null},{"paper":null,"title":"Simple is Better and Large is Not Enough: Towards Ensembling of Foundational Language Models","date":"2023-08-23","arxiv_id":"2308.12272","n_code_links":0,"syntology":null},{"paper":null,"title":"A Hybrid Machine Learning Model for Classifying Gene Mutations in Cancer using LSTM, BiLSTM, CNN, GRU, and GloVe","date":"2023-07-24","arxiv_id":"2307.14361","n_code_links":0,"syntology":null},{"paper":null,"title":"Resume Information Extraction via Post-OCR Text Processing","date":"2023-06-23","arxiv_id":"2306.13775","n_code_links":0,"syntology":null},{"paper":null,"title":"Research on Named Entity Recognition in Improved transformer with R-Drop structure","date":"2023-06-14","arxiv_id":"2306.08315","n_code_links":0,"syntology":null},{"paper":null,"title":"Domain-specific Continued Pretraining of Language Models for Capturing Long Context in Mental Health","date":"2023-04-20","arxiv_id":"2304.10447","n_code_links":0,"syntology":null},{"paper":null,"title":"Deep Learning for Opinion Mining and Topic Classification of Course Reviews","date":"2023-04-06","arxiv_id":"2304.03394","n_code_links":0,"syntology":null},{"paper":"/paper/trojtext-test-time-invisible-textual-trojan","title":"TrojText: Test-time Invisible Textual Trojan Insertion","date":"2023-03-03","arxiv_id":"2303.02242","n_code_links":1,"syntology":{"ran":3,"of":4,"unverified":1,"pointer_only":4}},{"paper":"/paper/hulat-at-semeval-2023-task-10-data","title":"HULAT at SemEval-2023 Task 10: Data augmentation for pre-trained transformers applied to the detection of sexism in social media","date":"2023-02-24","arxiv_id":"2302.12840","n_code_links":1,"syntology":null},{"paper":"/paper/a-benchmark-for-toxic-comment-classification","title":"A benchmark for toxic comment classification on Civil Comments dataset","date":"2023-01-26","arxiv_id":"2301.11125","n_code_links":1,"syntology":null},{"paper":"/paper/analyzing-semantic-faithfulness-of-language","title":"Analyzing Semantic Faithfulness of Language Models via Input Intervention on Question Answering","date":"2022-12-21","arxiv_id":"2212.10696","n_code_links":1,"syntology":null},{"paper":null,"title":"Learning-To-Embed: Adopting Transformer based models for E-commerce Products Representation Learning","date":"2022-12-07","arxiv_id":"2212.03725","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":23},{"task":"/task/language-modeling","name":"Language Modeling","papers":22},{"task":"/task/sentence","name":"Sentence","papers":20},{"task":"/task/question-answering","name":"Question Answering","papers":17},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":17},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":14},{"task":"/task/text-classification","name":"Text Classification","papers":13},{"task":"/task/reading-comprehension","name":"Reading Comprehension","papers":11},{"task":"/task/text-classification-1","name":"text-classification","papers":10},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":9},{"task":"/task/classification","name":"General Classification","papers":8},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":8},{"task":"/task/text-generation","name":"Text Generation","papers":8},{"task":"/task/named-entity-recognition-ner","name":"Named Entity Recognition (NER)","papers":7},{"task":"/task/representation-learning","name":"Representation Learning","papers":7},{"task":"/task/classification-1","name":"Classification","papers":6},{"task":"/task/named-entity-recognition-1","name":"Named Entity Recognition","papers":6},{"task":"/task/sentiment-classification","name":"Sentiment Classification","papers":6},{"task":"/task/translation","name":"Translation","papers":6},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":6}],"tasks_shown":20,"n_tasks":190,"usage_by_year":[{"year":"2019","papers":25},{"year":"2020","papers":58},{"year":"2021","papers":42},{"year":"2022","papers":14},{"year":"2023","papers":14},{"year":"2024","papers":11},{"year":"2025","papers":3}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/xlnet"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}