Methods › General › Stochastic Optimization › AdamW › Papers, page 3
AdamW
Papers archive 2025-07-28
archive papers tagged: 206 · with a code link: 114 · where Syntology ran a sample: 43 (38 with a run with no instrument failure, 5 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (43 of 206 tagged: 38 with a run with no instrument failure, 5 where every run was a failure of Syntology's instrument)
Page 3 of 3: papers 201 to 206 of 206, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Longformer: The Long-Document Transformer 10 Apr 2020 · 22 repositories · arXiv:2004.05150Syntology official (archive's flag): 9 ran · 22 ran (of which 4 constructed an object rather than computing a result; 14 with no instrument failure: 1 honoured, 1 violated, 12 with no contract checked; 8 where Syntology's instrument failed) · 13 unverified (of 35 harvested samples) · 5 pointer-only (licence)
-
Automated Pavement Crack Segmentation Using U-Net-based Convolutional Neural Network 7 Jan 2020 · 0 repositories · arXiv:2001.01912
-
Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks 27 May 2019 · 3 repositories · arXiv:1905.11286Syntology 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples) · 1 pointer-only (licence)
-
A unified theory of adaptive stochastic gradient descent as Bayesian filtering 1 May 2019 · 0 repositories
-
Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods 19 Jul 2018 · 1 repository · arXiv:1807.07540
-
Decoupled Weight Decay Regularization 14 Nov 2017 · 23 repositories · arXiv:1711.05101Syntology community repositories only · 18 ran (of which 6 constructed an object rather than computing a result; 18 with no instrument failure: 3 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 23 harvested samples) · 5 pointer-only (licence)