Methods › General › Stochastic Optimization › Adam

Adam

24,390 papers tagged archive 2025-07-28

archive 2025-07-28 Description, source and code snippet are the archive's method entry.

Adam is an adaptive learning rate optimization algorithm that utilises both momentum and scaling, combining the benefits of RMSProp and SGD w/th Momentum. The optimizer is designed to be appropriate for non-stationary objectives and problems with very noisy and/or sparse gradients.

The weight updates are performed as:

wₜ = wₜ₋₁ - ηm̂ₜ/(√(v̂ₜ) + ϵ)

with

m̂ₜ = mₜ/(1-βᵗ₁)

v̂ₜ = vₜ/(1-βᵗ₂)

mₜ = β₁mₜ₋₁ + (1-β₁)gₜ

vₜ = β₂vₜ₋₁ + (1-β₂)gₜ²

η is the step size/learning rate, around 1e-3 in the original paper. ϵ is a small number, typically 1e-8 or 1e-10, to prevent dividing by zero. β₁ and β₂ are forgetting parameters, with typical values 0.9 and 0.999, respectively.

Source: Adam: A Method for Stochastic OptimizationSee Code · pytorch/pytorch

Papers archive 2025-07-28

30 shown of 24,390, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.

Tasks archive 2025-07-28

20 shown of 2,551 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.

TaskPapers
Language Modelling2,917
Language Modeling2,286
Retrieval1,734
Question Answering1,420
RAG1,290
Sentence1,289
Decoder1,284
Retrieval-augmented Generation1,121
Translation1,009
Machine Translation905
Large Language Model811
Text Generation723
Semantic Segmentation707
Image Classification660
Transfer Learning651
Representation Learning596
Sentiment Analysis595
Object Detection585
Text Classification559
Classification554

Usage over time archive 2025-07-28

Papers per year tagged with Adam: 1986 to 2025, peak 6,749 6,749 0 1986: 1 paper 1986 1987: 0 papers 1988: 0 papers 1989: 0 papers 1990: 0 papers 1991: 0 papers 1991 1992: 0 papers 1993: 0 papers 1994: 0 papers 1995: 0 papers 1996: 0 papers 1996 1997: 0 papers 1998: 0 papers 1999: 0 papers 2000: 0 papers 2001: 0 papers 2001 2002: 0 papers 2003: 0 papers 2004: 0 papers 2005: 0 papers 2006: 0 papers 2006 2007: 0 papers 2008: 0 papers 2009: 0 papers 2010: 0 papers 2011: 0 papers 2011 2012: 0 papers 2013: 0 papers 2014: 1 paper 2015: 6 papers 2016: 23 papers 2016 2017: 70 papers 2018: 222 papers 2019: 1198 papers 2020: 2319 papers 2021: 3153 papers 2021 2022: 3169 papers 2023: 5041 papers 2024: 6749 papers 2025: 2438 papers 2025
Papers per year the archive tags with this method, by the paper's archive date (24,390 dated). Bars are counts, not a trend claim.

Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).

Categories archive 2025-07-28

Stochastic Optimization

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections