Methods › General › Stochastic Optimization › SGD › Papers, page 13
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 13 of 21: papers 1,201 to 1,300 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Attentional-Biased Stochastic Gradient Descent 13 Dec 2020 · 1 repository · arXiv:2012.06951
-
Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning Dynamics 8 Dec 2020 · 1 repository · arXiv:2012.04728Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Noise and Fluctuation of Finite Learning Rate Stochastic Gradient Descent 7 Dec 2020 · 0 repositories · arXiv:2012.03636
-
Effect of the initial configuration of weights on the training and function of artificial neural networks 4 Dec 2020 · 0 repositories · arXiv:2012.02550
-
Lookahead optimizer improves the performance of Convolutional Autoencoders for reconstruction of natural images 3 Dec 2020 · 0 repositories · arXiv:2012.05694
-
Stochastic Gradient Descent with Nonlinear Conjugate Gradient-Style Adaptive Momentum 3 Dec 2020 · 0 repositories · arXiv:2012.02188
-
Convergence and Sample Complexity of SGD in GANs 1 Dec 2020 · 0 repositories · arXiv:2012.00732
-
Curriculum Learning by Dynamic Instance Hardness 1 Dec 2020 · 0 repositories
-
On the universality of deep learning 1 Dec 2020 · 0 repositories
-
Robustness Analysis of Non-Convex Stochastic Gradient Descent using Biased Expectations 1 Dec 2020 · 0 repositories
-
Stochastic Gradient Descent in Correlated Settings: A Study on Gaussian Processes 1 Dec 2020 · 0 repositories
-
Towards Better Generalization of Adaptive Gradient Methods 1 Dec 2020 · 0 repositories
-
Eigenvalue-corrected Natural Gradient Based on a New Approximation 27 Nov 2020 · 0 repositories · arXiv:2011.13609
-
Adam^+: A Stochastic Method with Adaptive Variance Reduction 24 Nov 2020 · 0 repositories · arXiv:2011.11985
-
Convergence Analysis of Homotopy-SGD for non-convex optimization 20 Nov 2020 · 0 repositories · arXiv:2011.10298
-
Variational Laplace for Bayesian neural networks 20 Nov 2020 · 0 repositories · arXiv:2011.10443
-
Strong Data Augmentation Sanitizes Poisoning and Backdoor Attacks Without an Accuracy Tradeoff 18 Nov 2020 · 1 repository · arXiv:2011.09527
-
Contrastive Weight Regularization for Large Minibatch SGD 17 Nov 2020 · 0 repositories · arXiv:2011.08968
-
Avoiding Communication in Logistic Regression 16 Nov 2020 · 0 repositories · arXiv:2011.08281
-
Mixing ADAM and SGD: a Combined Optimization Method 16 Nov 2020 · 1 repository · arXiv:2011.08042
-
Ridge Rider: Finding Diverse Solutions by Following Eigenvectors of the Hessian 12 Nov 2020 · 0 repositories · arXiv:2011.06505
-
Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning Rate 4 Nov 2020 · 0 repositories · arXiv:2011.02538
-
Gradient-Based Empirical Risk Minimization using Local Polynomial Regression 4 Nov 2020 · 0 repositories · arXiv:2011.02522
-
Local SGD: Unified Theory and New Efficient Methods 3 Nov 2020 · 0 repositories · arXiv:2011.02828
-
SGB: Stochastic Gradient Bound Method for Optimizing Partition Functions 3 Nov 2020 · 0 repositories · arXiv:2011.01474
-
A Greedy Bit-flip Training Algorithm for Binarized Knowledge Graph Embeddings 1 Nov 2020 · 0 repositories
-
Hogwild! over Distributed Local Data Sets with Linearly Increasing Mini-Batch Sizes 27 Oct 2020 · 0 repositories · arXiv:2010.14763
-
Optimal Client Sampling for Federated Learning 26 Oct 2020 · 1 repository · arXiv:2010.13723
-
Distributed Saddle-Point Problems: Lower Bounds, Near-Optimal and Robust Algorithms 25 Oct 2020 · 0 repositories · arXiv:2010.13112
-
Inductive Bias of Gradient Descent for Weight Normalized Smooth Homogeneous Neural Nets 24 Oct 2020 · 1 repository · arXiv:2010.12909
-
Stochastic Gradient Descent Meets Distribution Regression 24 Oct 2020 · 0 repositories · arXiv:2010.12842
-
Adaptive Gradient Quantization for Data-Parallel SGD 23 Oct 2020 · 1 repository · arXiv:2010.12460Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Linearly Converging Error Compensated SGD 23 Oct 2020 · 1 repository · arXiv:2010.12292Syntology official: harvested, nothing ran · 0 ran · 3 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Computationally and Statistically Efficient Truncated Regression 22 Oct 2020 · 0 repositories · arXiv:2010.12000
-
Adaptive Gradient Method with Resilience and Momentum 21 Oct 2020 · 0 repositories · arXiv:2010.11041
-
How Data Augmentation affects Optimization for Linear Regression 21 Oct 2020 · 0 repositories · arXiv:2010.11171
-
A Flatter Loss for Bias Mitigation in Cross-dataset Facial Age Estimation 20 Oct 2020 · 0 repositories · arXiv:2010.10368
-
Dual Averaging is Surprisingly Effective for Deep Learning Optimization 20 Oct 2020 · 0 repositories · arXiv:2010.10502
-
Why Are Convolutional Nets More Sample-Efficient than Fully-Connected Nets? 16 Oct 2020 · 0 repositories · arXiv:2010.08515
-
AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed Gradients 15 Oct 2020 · 8 repositories · arXiv:2010.07468Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
RNN Training along Locally Optimal Trajectories via Frank-Wolfe Algorithm 12 Oct 2020 · 1 repository · arXiv:2010.05397
-
Towards Theoretically Understanding Why SGD Generalizes Better Than ADAM in Deep Learning 12 Oct 2020 · 0 repositories · arXiv:2010.05627
-
TaxoNN: A Light-Weight Accelerator for Deep Neural Network Training 11 Oct 2020 · 0 repositories · arXiv:2010.05197
-
AEGD: Adaptive Gradient Descent with Energy 10 Oct 2020 · 1 repository · arXiv:2010.05109
-
Double Forward Propagation for Memorized Batch Normalization 10 Oct 2020 · 0 repositories · arXiv:2010.04947
-
Theedhum Nandrum@Dravidian-CodeMix-FIRE2020: A Sentiment Polarity Classifier for YouTube Comments with Code-switching between Tamil, Malayalam and English 7 Oct 2020 · 1 repository · arXiv:2010.03189
-
Practical Precoding via Asynchronous Stochastic Successive Convex Approximation 3 Oct 2020 · 1 repository · arXiv:2010.01360
-
Variance-Reduced Methods for Machine Learning 2 Oct 2020 · 0 repositories · arXiv:2010.00892
-
Understanding Self-supervised Learning with Dual Deep Networks 1 Oct 2020 · 2 repositories · arXiv:2010.00578Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Momentum via Primal Averaging: Theoretical Insights and Learning Rate Schedules for Non-Convex Optimization 1 Oct 2020 · 1 repository · arXiv:2010.00406
-
Learning to Generate Image Source-Agnostic Universal Adversarial Perturbations 29 Sep 2020 · 0 repositories · arXiv:2009.13714
-
Apollo: An Adaptive Parameter-wise Diagonal Quasi-Newton Method for Nonconvex Stochastic Optimization 28 Sep 2020 · 3 repositories · arXiv:2009.13586Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Improved generalization by noise enhancement 28 Sep 2020 · 0 repositories · arXiv:2009.13094
-
Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal Effect 28 Sep 2020 · 2 repositories · arXiv:2009.12991Syntology official (archive's flag): 1 ran · 2 ran (of which 2 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 2 samples that ran constructed an object rather than computing a result (of 2 harvested samples) · 2 pointer-only (licence)
-
On Efficient Constructions of Checkpoints 28 Sep 2020 · 0 repositories · arXiv:2009.13003
-
Over-the-Air Federated Learning from Heterogeneous Data 27 Sep 2020 · 1 repository · arXiv:2009.12787
-
How Many Factors Influence Minima in SGD? 24 Sep 2020 · 0 repositories · arXiv:2009.11858
-
Anomalous diffusion dynamics of learning in deep neural networks 22 Sep 2020 · 1 repository · arXiv:2009.10588Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 11 harvested samples) · 4 pointer-only (licence)
-
Asynchronous Distributed Optimization with Stochastic Delays 22 Sep 2020 · 0 repositories · arXiv:2009.10717
-
Sparse Communication for Training Deep Networks 19 Sep 2020 · 0 repositories · arXiv:2009.09271
-
Low-Rank Training of Deep Neural Networks for Emerging Memory Technology 8 Sep 2020 · 0 repositories · arXiv:2009.03887
-
An Analysis of Alternating Direction Method of Multipliers for Feed-forward Neural Networks 6 Sep 2020 · 0 repositories · arXiv:2009.02825
-
HPSGD: Hierarchical Parallel SGD With Stale Gradients Featuring 6 Sep 2020 · 1 repository · arXiv:2009.02701
-
S-SGD: Symmetrical Stochastic Gradient Descent with Weight Noise Injection for Reaching Flat Minima 5 Sep 2020 · 0 repositories · arXiv:2009.02479
-
On Communication Compression for Distributed Optimization on Heterogeneous Data 4 Sep 2020 · 0 repositories · arXiv:2009.02388
-
Private Weighted Random Walk Stochastic Gradient Descent 3 Sep 2020 · 0 repositories · arXiv:2009.01790
-
Extreme Memorization via Scale of Initialization 31 Aug 2020 · 1 repository · arXiv:2008.13363
-
ROOT-SGD: Sharp Nonasymptotics and Near-Optimal Asymptotics in a Single Algorithm 28 Aug 2020 · 0 repositories · arXiv:2008.12690
-
A Fast and Robust BERT-based Dialogue State Tracker for Schema-Guided Dialogue Dataset 27 Aug 2020 · 1 repository · arXiv:2008.12335
-
Adversarially Robust Learning via Entropic Regularization 27 Aug 2020 · 0 repositories · arXiv:2008.12338
-
APMSqueeze: A Communication Efficient Adam-Preconditioned Momentum SGD Algorithm 26 Aug 2020 · 0 repositories · arXiv:2008.11343
-
PAGE: A Simple and Optimal Probabilistic Gradient Estimator for Nonconvex Optimization 25 Aug 2020 · 0 repositories · arXiv:2008.10898
-
Adaptive Serverless Learning 24 Aug 2020 · 0 repositories · arXiv:2008.10422
-
Periodic Stochastic Gradient Descent with Momentum for Decentralized Training 24 Aug 2020 · 0 repositories · arXiv:2008.10435
-
Surrogate NAS Benchmarks: Going Beyond the Limited Search Spaces of Tabular NAS Benchmarks 22 Aug 2020 · 1 repository · arXiv:2008.09777Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
A(DP)²SGD: Asynchronous Decentralized Parallel Stochastic Gradient Descent with Differential Privacy 21 Aug 2020 · 0 repositories · arXiv:2008.09246
-
Optimization of Graph Neural Networks with Natural Gradient Descent 21 Aug 2020 · 1 repository · arXiv:2008.09624
-
Obtaining Adjustable Regularization for Free via Iterate Averaging 15 Aug 2020 · 1 repository · arXiv:2008.06736Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Orthogonalized SGD and Nested Architectures for Anytime Neural Networks 15 Aug 2020 · 0 repositories · arXiv:2008.06635
-
Fast Dimension Independent Private AdaGrad on Publicly Estimated Subspaces 14 Aug 2020 · 0 repositories · arXiv:2008.06570
-
Privacy-Preserving Asynchronous Federated Learning Algorithms for Multi-Party Vertically Collaborative Learning 14 Aug 2020 · 0 repositories · arXiv:2008.06233
-
Three Variants of Differential Privacy: Lossless Conversion and Applications 14 Aug 2020 · 0 repositories · arXiv:2008.06529
-
Deep Networks with Fast Retraining 13 Aug 2020 · 0 repositories · arXiv:2008.07387
-
Holdout SGD: Byzantine Tolerant Federated Learning 11 Aug 2020 · 0 repositories · arXiv:2008.04612
-
An improved convergence analysis for decentralized online stochastic non-convex optimization 10 Aug 2020 · 0 repositories · arXiv:2008.04195
-
Mime: Mimicking Centralized Stochastic Algorithms in Federated Learning 8 Aug 2020 · 1 repository · arXiv:2008.03606Syntology 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
MLR-SNet: Transferable LR Schedules for Heterogeneous Tasks 29 Jul 2020 · 0 repositories · arXiv:2007.14546
-
A High Probability Analysis of Adaptive SGD with Momentum 28 Jul 2020 · 0 repositories · arXiv:2007.14294
-
Stochastic Normalized Gradient Descent with Momentum for Large-Batch Training 28 Jul 2020 · 0 repositories · arXiv:2007.13985
-
Multi-Level Local SGD for Heterogeneous Hierarchical Networks 27 Jul 2020 · 1 repository · arXiv:2007.13819Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 3 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 7 unverified (of 17 harvested samples) · 17 pointer-only (licence)
-
On the Regularization Effect of Stochastic Gradient Descent applied to Least Squares 27 Jul 2020 · 0 repositories · arXiv:2007.13288
-
CSER: Communication-efficient SGD with Error Reset 26 Jul 2020 · 0 repositories · arXiv:2007.13221
-
Train Like a (Var)Pro: Efficient Training of Neural Networks with Variable Projection 26 Jul 2020 · 1 repository · arXiv:2007.13171
-
Neural networks with late-phase weights 25 Jul 2020 · 2 repositories · arXiv:2007.12927Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
How to Democratise and Protect AI: Fair and Differentially Private Decentralised Deep Learning 18 Jul 2020 · 0 repositories · arXiv:2007.09370
-
On regularization of gradient descent, layer imbalance and flat minima 18 Jul 2020 · 0 repositories · arXiv:2007.09286
-
Distributed Reinforcement Learning of Targeted Grasping with Active Vision for Mobile Manipulators 16 Jul 2020 · 0 repositories · arXiv:2007.08082
-
Analysis of Q-learning with Adaptation and Momentum Restart for Gradient Descent 15 Jul 2020 · 0 repositories · arXiv:2007.07422
-
Gradient-based Hyperparameter Optimization Over Long Horizons 15 Jul 2020 · 1 repository · arXiv:2007.07869
-
Adaptive Periodic Averaging: A Practical Approach to Reducing Communication in Distributed Learning 13 Jul 2020 · 0 repositories · arXiv:2007.06134