Methods › General › Large Batch Optimization › AdaGrad › Papers, page 2
AdaGrad
Papers archive 2025-07-28
archive papers tagged: 191 · with a code link: 70 · where Syntology ran a sample: 23 (20 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (23 of 191 tagged: 20 with a run with no instrument failure, 3 where every run was a failure of Syntology's instrument)
Page 2 of 2: papers 101 to 191 of 191, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
L2M: Practical posterior Laplace approximation with optimization-driven second moment estimation 9 Jul 2021 · 0 repositories · arXiv:2107.04695
-
Private Adaptive Gradient Methods for Convex Optimization 25 Jun 2021 · 0 repositories · arXiv:2106.13756
-
A decreasing scaling transition scheme from Adam to SGD 12 Jun 2021 · 2 repositories · arXiv:2106.06749
-
Comparative Investigation of Learning Algorithms for Image Classification with Small Dataset 11 Jun 2021 · 0 repositories
-
Cross-Trajectory Representation Learning for Zero-Shot Generalization in RL 4 Jun 2021 · 1 repository · arXiv:2106.02193Syntology official (archive's flag): 7 ran · 7 ran (of which 2 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 11 pointer-only (licence)
-
Local Adaptivity in Federated Learning: Convergence and Consistency 4 Jun 2021 · 0 repositories · arXiv:2106.02305
-
Generalized AdaGrad (G-AdaGrad) and Adam: A State-Space Perspective 31 May 2021 · 0 repositories · arXiv:2106.00092
-
Learning to Relate Depth and Semantics for Unsupervised Domain Adaptation 17 May 2021 · 1 repository · arXiv:2105.07830Syntology official (archive's flag): 3 ran · 3 ran (of which 1 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
CTLR@WiC-TSV: Target Sense Verification using Marked Inputs andPre-trained Models 30 Apr 2021 · 0 repositories
-
ExplainaBoard: An Explainable Leaderboard for NLP 13 Apr 2021 · 1 repository · arXiv:2104.06387
-
Exploiting Adam-like Optimization Algorithms to Improve the Performance of Convolutional Neural Networks 26 Mar 2021 · 0 repositories · arXiv:2103.14689
-
Spatio-Temporal Neural Network for Fitting and Forecasting COVID-19 22 Mar 2021 · 0 repositories · arXiv:2103.11860
-
Categorical Foundations of Gradient-Based Learning 2 Mar 2021 · 0 repositories · arXiv:2103.01931
-
Variance Reduced Training with Stratified Sampling for Forecasting Models 2 Mar 2021 · 0 repositories · arXiv:2103.02062
-
SVRG Meets AdaGrad: Painless Variance Reduction 18 Feb 2021 · 0 repositories · arXiv:2102.09645
-
MetaGrad: Adaptation using Multiple Learning Rates in Online Learning 12 Feb 2021 · 0 repositories · arXiv:2102.06622
-
Adaptivity without Compromise: A Momentumized, Adaptive, Dual Averaged Gradient Method for Stochastic Optimization 26 Jan 2021 · 5 repositories · arXiv:2101.11075
-
Adam revisited: a weighted past gradients perspective 1 Jan 2021 · 0 repositories · arXiv:2101.00238
-
Adaptive Gradient Methods Can Be Provably Faster than SGD with Random Shuffling 1 Jan 2021 · 0 repositories
-
Convergent Adaptive Gradient Methods in Decentralized Optimization 1 Jan 2021 · 0 repositories
-
CADA: Communication-Adaptive Distributed Adam 31 Dec 2020 · 1 repository · arXiv:2012.15469Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Variance Reduction on General Adaptive Stochastic Mirror Descent 26 Dec 2020 · 0 repositories · arXiv:2012.13760
-
The Implicit Bias for Adaptive Optimization Algorithms on Homogeneous Neural Networks 11 Dec 2020 · 1 repository · arXiv:2012.06244
-
Asymptotic study of stochastic adaptive algorithm in non-convex landscape 10 Dec 2020 · 0 repositories · arXiv:2012.05640
-
Better Full-Matrix Regret via Parameter-Free Online Learning 1 Dec 2020 · 0 repositories
-
Towards Better Generalization of Adaptive Gradient Methods 1 Dec 2020 · 0 repositories
-
Sequential convergence of AdaGrad algorithm for smooth convex optimization 24 Nov 2020 · 0 repositories · arXiv:2011.12341
-
Which Minimizer Does My Neural Network Converge To? 4 Nov 2020 · 0 repositories · arXiv:2011.02408
-
Why Are Convolutional Nets More Sample-Efficient than Fully-Connected Nets? 16 Oct 2020 · 0 repositories · arXiv:2010.08515
-
Reparametrizing gradient descent 9 Oct 2020 · 0 repositories · arXiv:2010.04786
-
Fast Dimension Independent Private AdaGrad on Publicly Estimated Subspaces 14 Aug 2020 · 0 repositories · arXiv:2008.06570
-
Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs 5 Aug 2020 · 3 repositories · arXiv:2008.02312Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
A High Probability Analysis of Adaptive SGD with Momentum 28 Jul 2020 · 0 repositories · arXiv:2007.14294
-
Corner Proposal Network for Anchor-free, Two-stage Object Detection 27 Jul 2020 · 1 repository · arXiv:2007.13816
-
Adaptive Gradient Methods for Constrained Convex Optimization and Variational Inequalities 17 Jul 2020 · 0 repositories · arXiv:2007.08840
-
Adaptive Gradient Methods Can Be Provably Faster than SGD after Finite Epochs 12 Jun 2020 · 0 repositories · arXiv:2006.07037
-
Adaptive Gradient Methods Converge Faster with Over-Parameterization (but you should do a line-search) 11 Jun 2020 · 1 repository · arXiv:2006.06835Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
ADAHESSIAN: An Adaptive Second Order Optimizer for Machine Learning 1 Jun 2020 · 4 repositories · arXiv:2006.00719Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
A Simple Convergence Proof of Adam and Adagrad 5 Mar 2020 · 0 repositories · arXiv:2003.02395
-
Stagewise Enlargement of Batch Size for SGD-based Learning 26 Feb 2020 · 0 repositories · arXiv:2002.11601
-
Adaptive Online Learning with Varying Norms 10 Feb 2020 · 0 repositories · arXiv:2002.03963
-
Revisiting the Generalization of Adaptive Gradient Methods 1 Jan 2020 · 0 repositories
-
Towards Better Understanding of Adaptive Gradient Algorithms in Generative Adversarial Nets 26 Dec 2019 · 0 repositories · arXiv:1912.11940
-
Second-order Information in First-order Optimization Methods 20 Dec 2019 · 0 repositories · arXiv:1912.09926
-
Parameter Continuation Methods for the Optimization of Deep Neural Networks 16 Dec 2019 · 1 repository
-
Memory Efficient Adaptive Optimization 1 Dec 2019 · 1 repository
-
An Adaptive and Momental Bound Method for Stochastic Learning 27 Oct 2019 · 2 repositories · arXiv:1910.12249Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Implementation of a modified Nesterov's Accelerated quasi-Newton Method on Tensorflow 21 Oct 2019 · 0 repositories · arXiv:1910.09158
-
Adaptive Step Sizes in Variance Reduction via Regularization 15 Oct 2019 · 0 repositories · arXiv:1910.06532
-
diffGrad: An Optimization Method for Convolutional Neural Networks 12 Sep 2019 · 1 repository · arXiv:1909.11015
-
CTRL: A Conditional Transformer Language Model for Controllable Generation 11 Sep 2019 · 8 repositories · arXiv:1909.05858Syntology 14 ran (of which 3 constructed an object rather than computing a result; 11 with no instrument failure: 3 honoured, 1 violated, 7 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 14 harvested samples)
-
Meta-descent for Online, Continual Prediction 17 Jul 2019 · 0 repositories · arXiv:1907.07751
-
Augmenting Self-attention with Persistent Memory 2 Jul 2019 · 2 repositories · arXiv:1907.01470Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 2 honoured, 3 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Adaptively Preconditioned Stochastic Gradient Langevin Dynamics 10 Jun 2019 · 1 repository · arXiv:1906.04324
-
The Implicit Bias of AdaGrad on Separable Data 9 Jun 2019 · 1 repository · arXiv:1906.03559
-
AdaOja: Adaptive Learning Rates for Streaming PCA 28 May 2019 · 1 repository · arXiv:1905.12115
-
Hyper-Regularization: An Adaptive Choice for the Learning Rate in Gradient Descent 1 May 2019 · 0 repositories
-
Adaptive Gradient Methods with Dynamic Bound of Learning Rate 26 Feb 2019 · 5 repositories · arXiv:1902.09843
-
Global Convergence of Adaptive Gradient Methods for An Over-parameterized Neural Network 19 Feb 2019 · 0 repositories · arXiv:1902.07111
-
A Universal Algorithm for Variational Inequalities Adaptive to Smoothness and Noise 5 Feb 2019 · 0 repositories · arXiv:1902.01637
-
Compressing Gradient Optimizers via Count-Sketches 1 Feb 2019 · 1 repository · arXiv:1902.00179
-
Memory-Efficient Adaptive Optimization 30 Jan 2019 · 4 repositories · arXiv:1901.11150
-
A Sufficient Condition for Convergences of Adam and RMSProp 23 Nov 2018 · 0 repositories · arXiv:1811.09358
-
Practical Bayesian Learning of Neural Networks via Adaptive Optimisation Methods 8 Nov 2018 · 1 repository · arXiv:1811.03679
-
Riemannian Adaptive Optimization Methods 1 Oct 2018 · 1 repository · arXiv:1810.00760
-
Universal Stagewise Learning for Non-Convex Problems with Convergence on Averaged Solutions 20 Aug 2018 · 0 repositories · arXiv:1808.06296
-
A Unified Analysis of AdaGrad with Weighted Aggregation and Momentum Acceleration 10 Aug 2018 · 0 repositories · arXiv:1808.03408
-
On the Convergence of A Class of Adam-Type Algorithms for Non-Convex Optimization 8 Aug 2018 · 0 repositories · arXiv:1808.02941
-
SADAGRAD: Strongly Adaptive Stochastic Gradient Methods 1 Jul 2018 · 0 repositories
-
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes 5 Jun 2018 · 1 repository · arXiv:1806.01811
-
On the Convergence of Stochastic Gradient Descent with Adaptive Stepsizes 21 May 2018 · 0 repositories · arXiv:1805.08114
-
Block Mean Approximation for Efficient Second Order Optimization 16 Apr 2018 · 0 repositories · arXiv:1804.05484
-
Shampoo: Preconditioned Stochastic Tensor Optimization 26 Feb 2018 · 3 repositories · arXiv:1802.09568Syntology 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
LSH-SAMPLING BREAKS THE COMPUTATIONAL CHICKEN-AND-EGG LOOP IN ADAPTIVE STOCHASTIC GRADIENT ESTIMATION 1 Jan 2018 · 0 repositories
-
Improving Generalization Performance by Switching from Adam to SGD 20 Dec 2017 · 6 repositories · arXiv:1712.07628
-
AdaBatch: Efficient Gradient Aggregation Rules for Sequential and Parallel Stochastic Gradient Methods 6 Nov 2017 · 0 repositories · arXiv:1711.01761
-
Why ADAGRAD Fails for Online Topic Modeling 1 Sep 2017 · 0 repositories
-
A Unified Approach to Adaptive Regularization in Online and Stochastic Optimization 20 Jun 2017 · 0 repositories · arXiv:1706.06569
-
YellowFin and the Art of Momentum Tuning 12 Jun 2017 · 2 repositories · arXiv:1706.03471Syntology 0 ran · 1 unverified (of 1 harvested sample)
-
The Marginal Value of Adaptive Gradient Methods in Machine Learning 23 May 2017 · 3 repositories · arXiv:1705.08292
-
Efficient Parallel Translating Embedding For Knowledge Graphs 30 Mar 2017 · 1 repository · arXiv:1703.10316
-
Wasserstein GAN 26 Jan 2017 · 120 repositories · arXiv:1701.07875Syntology 16 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 1 honoured, 0 violated, 3 with no contract checked; 12 where Syntology's instrument failed) · 6 unverified (of 22 harvested samples) · 13 pointer-only (licence)
-
Improving Neural Language Models with a Continuous Cache 13 Dec 2016 · 14 repositories · arXiv:1612.04426
-
Scalable Adaptive Stochastic Optimization Using Random Projections 21 Nov 2016 · 0 repositories · arXiv:1611.06652
-
Relativistic Monte Carlo 14 Sep 2016 · 0 repositories · arXiv:1609.04388
-
CompAdaGrad: A Compressed, Complementary, Computationally-Efficient Adaptive Gradient Method 12 Sep 2016 · 0 repositories · arXiv:1609.03319
-
Bridging the Gap between Stochastic Gradient MCMC and Stochastic Optimization 25 Dec 2015 · 1 repository · arXiv:1512.07962
-
Speed learning on the fly 8 Nov 2015 · 0 repositories · arXiv:1511.02540
-
adaQN: An Adaptive Quasi-Newton Algorithm for Training RNNs 4 Nov 2015 · 0 repositories · arXiv:1511.01169
-
Path-SGD: Path-Normalized Optimization in Deep Neural Networks 8 Jun 2015 · 1 repository · arXiv:1506.02617
-
Dropout Training as Adaptive Regularization 4 Jul 2013 · 0 repositories · arXiv:1307.1493