{"url":"/method/albert","slug":"albert","name":"ALBERT","full_name":"ALBERT","full_name_withheld":false,"description_markdown":"**ALBERT** is a [Transformer](https://paperswithcode.com/method/transformer) architecture based on [BERT](https://paperswithcode.com/method/bert) but with much fewer parameters. It achieves this through two parameter reduction techniques. The first is a factorized embeddings parameterization. By decomposing the large vocabulary embedding matrix into two small matrices, the size of the hidden layers is separated from the size of vocabulary embedding. This makes it easier to grow the hidden size without significantly increasing the parameter size of the vocabulary embeddings. The second technique is cross-layer parameter sharing. This technique prevents the parameter from growing with the depth of the network. \r\n\r\nAdditionally, ALBERT utilises a self-supervised loss for sentence-order prediction (SOP). SOP primary focuses on inter-sentence coherence and is designed to address the ineffectiveness of the next sentence prediction (NSP) loss proposed in the original BERT.","description_state":"present","introduced_year":null,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":"https://arxiv.org/abs/1909.11942v6","title":"ALBERT: A Lite BERT for Self-supervised Learning of Language Representations","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Natural Language Processing","area_id":"natural-language-processing","collection":"Transformers","url":"/methods/category/transformers","pwc_aliases":[]}],"n_papers_tagged":172,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"ALBERT: Advanced Localization and Bidirectional Encoder Representations from Transformers for Automotive Damage Evaluation","date":"2025-06-12","arxiv_id":"2506.10524","n_code_links":0,"syntology":null},{"paper":null,"title":"Rapid yet accurate Tile-circuit and device modeling for Analog In-Memory Computing","date":"2025-05-05","arxiv_id":"2506.00004","n_code_links":0,"syntology":null},{"paper":"/paper/don-t-fight-hallucinations-use-them","title":"Don't Fight Hallucinations, Use Them: Estimating Image Realism using NLI over Atomic Facts","date":"2025-03-20","arxiv_id":"2503.15948","n_code_links":1,"syntology":null},{"paper":null,"title":"Efficient or Powerful? Trade-offs Between Machine Learning and Deep Learning for Mental Illness Detection on Social Media","date":"2025-03-03","arxiv_id":"2503.01082","n_code_links":0,"syntology":null},{"paper":"/paper/robust-bias-detection-in-mlms-and-its","title":"Robust Bias Detection in MLMs and its Application to Human Trait Ratings","date":"2025-02-21","arxiv_id":"2502.15600","n_code_links":1,"syntology":null},{"paper":null,"title":"Meursault as a Data Point","date":"2025-02-03","arxiv_id":"2502.01364","n_code_links":0,"syntology":null},{"paper":"/paper/punctuation-s-semantic-role-between-brain-and","title":"Aligning Brain Activity with Advanced Transformer Models: Exploring the Role of Punctuation in Semantic Processing","date":"2025-01-10","arxiv_id":"2501.06278","n_code_links":1,"syntology":null},{"paper":"/paper/a-comparative-analysis-of-transformer-and","title":"A Comparative Analysis of Transformer and LSTM Models for Detecting Suicidal Ideation on Reddit","date":"2024-11-23","arxiv_id":"2411.15404","n_code_links":1,"syntology":null},{"paper":"/paper/bert-based-approach-for-automating-course","title":"BERT-Based Approach for Automating Course Articulation Matrix Construction with Explainable AI","date":"2024-11-21","arxiv_id":"2411.14254","n_code_links":1,"syntology":null},{"paper":"/paper/protransformer-robustify-transformers-via","title":"ProTransformer: Robustify Transformers via Plug-and-Play Paradigm","date":"2024-10-30","arxiv_id":"2410.23182","n_code_links":1,"syntology":null},{"paper":null,"title":"A Bayesian Perspective on the Maximum Score Problem","date":"2024-10-22","arxiv_id":"2410.17153","n_code_links":0,"syntology":null},{"paper":null,"title":"Meta-RTL: Reinforcement-Based Meta-Transfer Learning for Low-Resource Commonsense Reasoning","date":"2024-09-27","arxiv_id":"2409.19075","n_code_links":0,"syntology":null},{"paper":"/paper/profiling-patient-transcript-using-large","title":"Profiling Patient Transcript Using Large Language Model Reasoning Augmentation for Alzheimer's Disease Detection","date":"2024-09-19","arxiv_id":"2409.12541","n_code_links":1,"syntology":null},{"paper":null,"title":"BioMNER: A Dataset for Biomedical Method Entity Recognition","date":"2024-06-28","arxiv_id":"2406.20038","n_code_links":0,"syntology":null},{"paper":null,"title":"Concept Formation and Alignment in Language Models: Bridging Statistical Patterns in Latent Space to Concept Taxonomy","date":"2024-06-08","arxiv_id":"2406.05315","n_code_links":0,"syntology":null},{"paper":null,"title":"Effect of antibody levels on the spread of disease in multiple infections","date":"2024-05-31","arxiv_id":"2405.20702","n_code_links":0,"syntology":null},{"paper":"/paper/ceebert-cross-domain-inference-in-early-exit","title":"CEEBERT: Cross-Domain Inference in Early Exit BERT","date":"2024-05-23","arxiv_id":"2405.15039","n_code_links":1,"syntology":{"ran":2,"of":3,"unverified":1,"pointer_only":3}},{"paper":null,"title":"A Named Entity Recognition and Topic Modeling-based Solution for Locating and Better Assessment of Natural Disasters in Social Media","date":"2024-05-01","arxiv_id":"2405.00903","n_code_links":0,"syntology":null},{"paper":null,"title":"Exploring Internal Numeracy in Language Models: A Case Study on ALBERT","date":"2024-04-25","arxiv_id":"2404.16574","n_code_links":0,"syntology":null},{"paper":"/paper/evaluating-subword-tokenization-alien-subword","title":"Evaluating Subword Tokenization: Alien Subword Composition and OOV Generalization Challenge","date":"2024-04-20","arxiv_id":"2404.13292","n_code_links":1,"syntology":null},{"paper":null,"title":"mALBERT: Is a Compact Multilingual BERT Model Still Worth It?","date":"2024-03-27","arxiv_id":"2403.18338","n_code_links":0,"syntology":null},{"paper":"/paper/an-exploratory-study-on-automatic","title":"An Exploratory Study on Automatic Identification of Assumptions in the Development of Deep Learning Frameworks","date":"2024-01-08","arxiv_id":"2401.03653","n_code_links":1,"syntology":null},{"paper":"/paper/tensor-aware-energy-accounting","title":"Tensor-Aware Energy Accounting","date":"2023-11-19","arxiv_id":"2311.11424","n_code_links":1,"syntology":null},{"paper":null,"title":"Generative AI for Hate Speech Detection: Evaluation and Findings","date":"2023-11-16","arxiv_id":"2311.09993","n_code_links":0,"syntology":null},{"paper":"/paper/amplify-attention-based-mixup-for-performance","title":"AMPLIFY:Attention-based Mixup for Performance Improvement and Label Smoothing in Transformer","date":"2023-09-22","arxiv_id":"2309.12689","n_code_links":1,"syntology":null},{"paper":"/paper/identification-of-the-relevance-of-comments","title":"Identification of the Relevance of Comments in Codes Using Bag of Words and Transformer Based Models","date":"2023-08-11","arxiv_id":"2308.06144","n_code_links":1,"syntology":null},{"paper":"/paper/performance-analysis-of-transformer-based","title":"Performance Analysis of Transformer Based Models (BERT, ALBERT and RoBERTa) in Fake News Detection","date":"2023-08-09","arxiv_id":"2308.04950","n_code_links":1,"syntology":null},{"paper":null,"title":"Gradient-Based Word Substitution for Obstinate Adversarial Examples Generation in Language Models","date":"2023-07-24","arxiv_id":"2307.12507","n_code_links":0,"syntology":null},{"paper":null,"title":"SparseOptimizer: Sparsify Language Models through Moreau-Yosida Regularization and Accelerate via Compiler Co-design","date":"2023-06-27","arxiv_id":"2306.15656","n_code_links":0,"syntology":null},{"paper":null,"title":"F-PABEE: Flexible-patience-based Early Exiting for Single-label and Multi-label text Classification Tasks","date":"2023-05-21","arxiv_id":"2305.11916","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":29},{"task":"/task/language-modeling","name":"Language Modeling","papers":24},{"task":"/task/sentence","name":"Sentence","papers":23},{"task":"/task/text-classification","name":"Text Classification","papers":15},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":14},{"task":"/task/question-answering","name":"Question Answering","papers":13},{"task":"/task/text-classification-1","name":"text-classification","papers":13},{"task":"/task/named-entity-recognition-ner","name":"Named Entity Recognition (NER)","papers":10},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":10},{"task":"/task/named-entity-recognition-1","name":"Named Entity Recognition","papers":9},{"task":"/task/cg","name":"NER","papers":8},{"task":"/task/named-entity-recognition","name":"named-entity-recognition","papers":8},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":7},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":7},{"task":"/task/reading-comprehension","name":"Reading Comprehension","papers":7},{"task":"/task/classification","name":"General Classification","papers":6},{"task":"/task/machine-reading-comprehension","name":"Machine Reading Comprehension","papers":6},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":6},{"task":"/task/word-embeddings","name":"Word Embeddings","papers":6},{"task":"/task/classification-1","name":"Classification","papers":5}],"tasks_shown":20,"n_tasks":167,"usage_by_year":[{"year":"2019","papers":2},{"year":"2020","papers":46},{"year":"2021","papers":63},{"year":"2022","papers":26},{"year":"2023","papers":13},{"year":"2024","papers":15},{"year":"2025","papers":7}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/albert"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}