{"url":"/method/adafactor","slug":"adafactor","name":"Adafactor","full_name":"Adafactor","full_name_withheld":false,"description_markdown":"**Adafactor** is a stochastic optimization method based on [Adam](https://paperswithcode.com/method/adam) that reduces memory usage while retaining the empirical benefits of adaptivity. This is achieved through maintaining a factored representation of the squared gradient accumulator across training steps. Specifically, by tracking moving averages of the row and column sums of the squared gradients for matrix-valued variables, we are able to reconstruct a low-rank approximation of the exponentially smoothed accumulator at each training step that is optimal with respect to the generalized Kullback-Leibler divergence. For an $n \\times m$ matrix, this reduces the memory requirements from $O(n m)$ to $O(n + m)$. \r\n\r\nInstead of defining the optimization algorithm in terms of absolute step sizes {$\\alpha_t$}$\\_{t=1}^T$, the authors define the optimization algorithm in terms of relative step sizes {$\\rho_t$}$\\_{t=1}^T$, which get multiplied by the scale of the parameters. The scale of a parameter vector or matrix is defined as the root-mean-square of its components, lower-bounded by a small constant $\\epsilon_2$.  The reason for this lower bound is to allow zero-initialized parameters to escape 0. \r\n\r\nProposed hyperparameters are: $\\epsilon\\_{1} = 10^{-30}$, $\\epsilon\\_{2} = 10^{-3}$, $d=1$, $p\\_{t} = \\min\\left(10^{-2}, \\frac{1}{\\sqrt{t}}\\right)$, $\\hat{\\beta}\\_{2\\_{t}} = 1 - t^{-0.8}$.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Adafactor: Adaptive Learning Rates with Sublinear Memory Cost","paper":"/paper/adafactor-adaptive-learning-rates-with","first_author":"Noam Shazeer","n_authors":2,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/adafactor-adaptive-learning-rates-with"},"source":{"url":"http://arxiv.org/abs/1804.04235v1","title":"Adafactor: Adaptive Learning Rates with Sublinear Memory Cost","url_on_a_paper_host":true},"code_snippet_url":"https://github.com/DeadAt0m/adafactor-pytorch/blob/561e627239c29c0be11256171a795b49e0404098/adafactor.py#L7","code_snippet_url_on_a_code_host":true,"categories":[{"area":"General","area_id":"general","collection":"Large Batch Optimization","url":"/methods/category/large-batch-optimization","pwc_aliases":[]},{"area":"General","area_id":"general","collection":"Stochastic Optimization","url":"/methods/category/stochastic-optimization","pwc_aliases":[]}],"n_papers_tagged":733,"archive_num_papers":733,"papers_newest_first":[{"paper":null,"title":"Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems","date":"2025-07-08","arxiv_id":"2507.05940","n_code_links":0,"syntology":null},{"paper":null,"title":"I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution","date":"2025-06-18","arxiv_id":"2506.17323","n_code_links":0,"syntology":null},{"paper":null,"title":"Fretting-Transformer: Encoder-Decoder Model for MIDI to Tablature Transcription","date":"2025-06-17","arxiv_id":"2506.14223","n_code_links":0,"syntology":null},{"paper":null,"title":"A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation","date":"2025-06-09","arxiv_id":"2506.08210","n_code_links":0,"syntology":null},{"paper":null,"title":"A Multi-Dataset Evaluation of Models for Automated Vulnerability Repair","date":"2025-06-05","arxiv_id":"2506.04987","n_code_links":0,"syntology":null},{"paper":null,"title":"Decom-Renorm-Merge: Model Merging on the Right Space Improves Multitasking","date":"2025-05-29","arxiv_id":"2505.23117","n_code_links":0,"syntology":null},{"paper":"/paper/shioenv-a-cli-behavior-capturing-environment","title":"ShIOEnv: A CLI Behavior-Capturing Environment Enabling Grammar-Guided Command Synthesis for Dataset Curation","date":"2025-05-23","arxiv_id":"2505.18374","n_code_links":1,"syntology":null},{"paper":"/paper/logicase-effective-test-case-generation-from","title":"LogiCase: Effective Test Case Generation from Logical Description in Competitive Programming","date":"2025-05-21","arxiv_id":"2505.15039","n_code_links":0,"syntology":{"ran":11,"of":19,"unverified":8,"pointer_only":19}},{"paper":"/paper/eeg-to-text-translation-a-model-for","title":"EEG-to-Text Translation: A Model for Deciphering Human Brain Activity","date":"2025-05-20","arxiv_id":"2505.13936","n_code_links":1,"syntology":null},{"paper":"/paper/masking-in-multi-hop-qa-an-analysis-of-how","title":"Masking in Multi-hop QA: An Analysis of How Language Models Perform with Context Permutation","date":"2025-05-16","arxiv_id":"2505.11754","n_code_links":1,"syntology":null},{"paper":null,"title":"Multilingual Machine Translation with Quantum Encoder Decoder Attention-based Convolutional Variational Circuits","date":"2025-05-14","arxiv_id":"2505.09407","n_code_links":0,"syntology":null},{"paper":"/paper/cardioformer-advancing-ai-in-ecg-analysis","title":"Cardioformer: Advancing AI in ECG Analysis with Multi-Granularity Patching and ResNet","date":"2025-05-08","arxiv_id":"2505.05538","n_code_links":1,"syntology":null},{"paper":null,"title":"Performance Evaluation of Large Language Models in Bangla Consumer Health Query Summarization","date":"2025-05-08","arxiv_id":"2505.05070","n_code_links":0,"syntology":null},{"paper":"/paper/gascade-grouped-summarization-of-adverse-drug","title":"GASCADE: Grouped Summarization of Adverse Drug Event for Enhanced Cancer Pharmacovigilance","date":"2025-05-07","arxiv_id":"2505.04284","n_code_links":1,"syntology":null},{"paper":null,"title":"A review of DNA restriction-free overlapping sequence cloning techniques for synthetic biology","date":"2025-05-06","arxiv_id":"2505.03681","n_code_links":0,"syntology":null},{"paper":null,"title":"JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry","date":"2025-04-29","arxiv_id":"2504.20849","n_code_links":0,"syntology":null},{"paper":null,"title":"Large Language Models are Qualified Benchmark Builders: Rebuilding Pre-Training Datasets for Advancing Code Intelligence Tasks","date":"2025-04-28","arxiv_id":"2504.19444","n_code_links":0,"syntology":null},{"paper":null,"title":"Fast-Powerformer: A Memory-Efficient Transformer for Accurate Mid-Term Wind Power Forecasting","date":"2025-04-15","arxiv_id":"2504.10923","n_code_links":0,"syntology":null},{"paper":"/paper/sigma-a-dataset-for-text-to-code-semantic","title":"Sigma: A dataset for text-to-code semantic parsing with statistical analysis","date":"2025-04-05","arxiv_id":"2504.04301","n_code_links":1,"syntology":null},{"paper":null,"title":"Advancing Sentiment Analysis in Tamil-English Code-Mixed Texts: Challenges and Transformer-Based Solutions","date":"2025-03-30","arxiv_id":"2503.23295","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Knowledge Graph Completion with Entity Neighborhood and Relation Context","date":"2025-03-29","arxiv_id":"2503.23205","n_code_links":0,"syntology":null},{"paper":"/paper/scaling-down-text-encoders-of-text-to-image","title":"Scaling Down Text Encoders of Text-to-Image Diffusion Models","date":"2025-03-25","arxiv_id":"2503.19897","n_code_links":1,"syntology":null},{"paper":"/paper/exploring-training-and-inference-scaling-laws","title":"Exploring Training and Inference Scaling Laws in Generative Retrieval","date":"2025-03-24","arxiv_id":"2503.18941","n_code_links":1,"syntology":null},{"paper":null,"title":"Predicting the Road Ahead: A Knowledge Graph based Foundation Model for Scene Understanding in Autonomous Driving","date":"2025-03-24","arxiv_id":"2503.18730","n_code_links":0,"syntology":null},{"paper":null,"title":"Enhancing Code LLM Training with Programmer Attention","date":"2025-03-19","arxiv_id":"2503.14936","n_code_links":0,"syntology":null},{"paper":null,"title":"DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models","date":"2025-03-17","arxiv_id":"2503.12885","n_code_links":0,"syntology":null},{"paper":null,"title":"ARLED: Leveraging LED-based ARMAN Model for Abstractive Summarization of Persian Long Documents","date":"2025-03-13","arxiv_id":"2503.10233","n_code_links":0,"syntology":null},{"paper":null,"title":"A LongFormer-Based Framework for Accurate and Efficient Medical Text Summarization","date":"2025-03-10","arxiv_id":"2503.06888","n_code_links":0,"syntology":null},{"paper":"/paper/roamify-designing-and-evaluating-an-llm-based","title":"Roamify: Designing and Evaluating an LLM Based Google Chrome Extension for Personalised Itinerary Planning","date":"2025-03-10","arxiv_id":"2504.10489","n_code_links":1,"syntology":null},{"paper":null,"title":"Seeing Delta Parameters as JPEG Images: Data-Free Delta Compression with Discrete Cosine Transform","date":"2025-03-09","arxiv_id":"2503.06676","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/language-modelling","name":"Language Modelling","papers":128},{"task":"/task/language-modeling","name":"Language Modeling","papers":100},{"task":"/task/question-answering","name":"Question Answering","papers":86},{"task":"/task/decoder","name":"Decoder","papers":79},{"task":"/task/text-generation","name":"Text Generation","papers":65},{"task":"/task/sentence","name":"Sentence","papers":58},{"task":"/task/translation","name":"Translation","papers":47},{"task":"/task/machine-translation","name":"Machine Translation","papers":41},{"task":"/task/retrieval","name":"Retrieval","papers":40},{"task":"/task/transfer-learning","name":"Transfer Learning","papers":40},{"task":"/task/natural-language-understanding","name":"Natural Language Understanding","papers":29},{"task":"/task/abstractive-text-summarization","name":"Abstractive Text Summarization","papers":23},{"task":"/task/semantic-parsing","name":"Semantic Parsing","papers":23},{"task":"/task/sentiment-analysis","name":"Sentiment Analysis","papers":22},{"task":"/task/natural-language-inference","name":"Natural Language Inference","papers":21},{"task":"/task/code-generation","name":"Code Generation","papers":20},{"task":"/task/data-augmentation","name":"Data Augmentation","papers":20},{"task":"/task/text-summarization","name":"Text Summarization","papers":19},{"task":"/task/diversity","name":"Diversity","papers":18},{"task":"/task/knowledge-graphs","name":"Knowledge Graphs","papers":17}],"tasks_shown":20,"n_tasks":492,"usage_by_year":[{"year":"2018","papers":1},{"year":"2019","papers":2},{"year":"2020","papers":37},{"year":"2021","papers":112},{"year":"2022","papers":168},{"year":"2023","papers":202},{"year":"2024","papers":158},{"year":"2025","papers":53}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/adafactor"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}