Methods › General › Stochastic Optimization › SGD › Papers, page 14
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 14 of 21: papers 1,301 to 1,400 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Quantitative Propagation of Chaos for SGD in Wide Neural Networks 13 Jul 2020 · 0 repositories · arXiv:2007.06352
-
AdaScale SGD: A User-Friendly Algorithm for Distributed Training 9 Jul 2020 · 1 repository · arXiv:2007.05105Syntology 0 ran · 2 unverified (of 2 harvested samples)
-
How benign is benign overfitting? 8 Jul 2020 · 0 repositories · arXiv:2007.04028
-
Bypassing the Ambient Dimension: Private SGD with Gradient Subspace Identification 7 Jul 2020 · 0 repositories · arXiv:2007.03813
-
Ridge Regression with Over-Parametrized Two-Layer Networks Converge to Ridgelet Spectrum 7 Jul 2020 · 0 repositories · arXiv:2007.03441
-
ResRep: Lossless CNN Pruning via Decoupling Remembering and Forgetting 7 Jul 2020 · 6 repositories · arXiv:2007.03260
-
Streaming Complexity of SVMs 7 Jul 2020 · 0 repositories · arXiv:2007.03633
-
Understanding the Impact of Model Incoherence on Convergence of Incremental SGD with Random Reshuffle 7 Jul 2020 · 0 repositories · arXiv:2007.03509
-
TDprop: Does Jacobi Preconditioning Help Temporal Difference Learning? 6 Jul 2020 · 0 repositories · arXiv:2007.02786Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Weak error analysis for stochastic gradient descent optimization algorithms 3 Jul 2020 · 0 repositories · arXiv:2007.02723
-
Adaptive Braking for Mitigating Gradient Delay 2 Jul 2020 · 0 repositories · arXiv:2007.01397
-
Balancing Rates and Variance via Adaptive Batch-Size for Stochastic Optimization Problems 2 Jul 2020 · 0 repositories · arXiv:2007.01219
-
ECPE-2D: Emotion-Cause Pair Extraction based on Joint Two-Dimensional Representation, Interaction and Prediction 1 Jul 2020 · 1 repository
-
Online Robust Regression via SGD on the l1 loss 1 Jul 2020 · 0 repositories · arXiv:2007.00399
-
AdaSGD: Bridging the gap between SGD and Adam 30 Jun 2020 · 0 repositories · arXiv:2006.16541
-
Adai: Separating the Effects of Adaptive Learning Rate and Momentum Inertia 29 Jun 2020 · 1 repository · arXiv:2006.15815Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 3 pointer-only (licence)
-
Understanding Gradient Clipping in Private SGD: A Geometric Perspective 27 Jun 2020 · 0 repositories · arXiv:2006.15429
-
Is SGD a Bayesian sampler? Well, almost 26 Jun 2020 · 0 repositories · arXiv:2006.15191
-
On the Generalization Benefit of Noise in Stochastic Gradient Descent 26 Jun 2020 · 0 repositories · arXiv:2006.15081
-
Stability Enhanced Privacy and Applications in Private Stochastic Gradient Descent 25 Jun 2020 · 0 repositories · arXiv:2006.14360
-
Dynamic of Stochastic Gradient Descent with State-Dependent Noise 24 Jun 2020 · 0 repositories · arXiv:2006.13719
-
Private Stochastic Non-Convex Optimization: Adaptive Algorithms and Tighter Generalization Bounds 24 Jun 2020 · 0 repositories · arXiv:2006.13501
-
Spherical Perspective on Learning with Normalization Layers 23 Jun 2020 · 1 repository · arXiv:2006.13382
-
Byzantine-Resilient High-Dimensional Federated Learning 22 Jun 2020 · 0 repositories · arXiv:2006.13041
-
MaxVA: Fast Adaptation of Step Sizes by Maximizing Observed Variance of Gradients 21 Jun 2020 · 1 repository · arXiv:2006.11918
-
Generalisation Guarantees for Continual Learning with Orthogonal Gradient Descent 21 Jun 2020 · 1 repository · arXiv:2006.11942Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
How do SGD hyperparameters in natural training affect adversarial robustness? 20 Jun 2020 · 0 repositories · arXiv:2006.11604
-
Training (Overparametrized) Neural Networks in Near-Linear Time 20 Jun 2020 · 0 repositories · arXiv:2006.11648
-
Unified Analysis of Stochastic Gradient Methods for Composite Convex and Smooth Optimization 20 Jun 2020 · 0 repositories · arXiv:2006.11573
-
DEED: A General Quantization Scheme for Communication Efficiency in Bits 19 Jun 2020 · 0 repositories · arXiv:2006.11401
-
Differentially Private Variational Autoencoders with Term-wise Gradient Aggregation 19 Jun 2020 · 0 repositories · arXiv:2006.11204
-
On the Almost Sure Convergence of Stochastic Gradient Descent in Non-Convex Problems 19 Jun 2020 · 0 repositories · arXiv:2006.11144
-
SGD for Structured Nonconvex Functions: Learning Rates, Minibatching and Interpolation 18 Jun 2020 · 0 repositories · arXiv:2006.10311
-
Stochastic Gradient Descent in Hilbert Scales: Smoothness, Preconditioning and Earlier Stopping 18 Jun 2020 · 0 repositories · arXiv:2006.10840
-
Communication-Efficient Robust Federated Learning Over Heterogeneous Datasets 17 Jun 2020 · 0 repositories · arXiv:2006.09992
-
Learning Rates as a Function of Batch Size: A Random Matrix Theory Approach to Neural Network Training 16 Jun 2020 · 1 repository · arXiv:2006.09092
-
Directional Pruning of Deep Neural Networks 16 Jun 2020 · 1 repository · arXiv:2006.09358Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Federated Accelerated Stochastic Gradient Descent 16 Jun 2020 · 1 repository · arXiv:2006.08950
-
Flatness is a False Friend 16 Jun 2020 · 0 repositories · arXiv:2006.09091
-
Hausdorff Dimension, Heavy Tails, and Generalization in Neural Networks 16 Jun 2020 · 1 repository · arXiv:2006.09313Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 12 harvested samples)
-
On sparse connectivity, adversarial robustness, and a novel model of the artificial neuron 16 Jun 2020 · 0 repositories · arXiv:2006.09510
-
Fine-Grained Analysis of Stability and Generalization for Stochastic Gradient Descent 15 Jun 2020 · 0 repositories · arXiv:2006.08157
-
Shape Matters: Understanding the Implicit Bias of the Noise Covariance 15 Jun 2020 · 1 repository · arXiv:2006.08680
-
AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant Weights 15 Jun 2020 · 4 repositories · arXiv:2006.08217Syntology official: no sample here; runs from other or unrecorded repositories · 4 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Spherical Motion Dynamics: Learning Dynamics of Neural Network with Normalization, Weight Decay, and SGD 15 Jun 2020 · 0 repositories · arXiv:2006.08419
-
An Analysis of Constant Step Size SGD in the Non-convex Regime: Asymptotic Normality and Bias 14 Jun 2020 · 0 repositories · arXiv:2006.07904
-
Topology-aware Differential Privacy for Decentralized Image Classification 14 Jun 2020 · 0 repositories · arXiv:2006.07817
-
Almost sure convergence rates for Stochastic Gradient Descent and Stochastic Heavy Ball 14 Jun 2020 · 0 repositories · arXiv:2006.07867
-
Auditing Differentially Private Machine Learning: How Private is Private SGD? 13 Jun 2020 · 1 repository · arXiv:2006.07709Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
The Pitfalls of Simplicity Bias in Neural Networks 13 Jun 2020 · 2 repositories · arXiv:2006.07710
-
A Unified Analysis of Stochastic Gradient Methods for Nonconvex Federated Optimization 12 Jun 2020 · 0 repositories · arXiv:2006.07013
-
Adaptive Gradient Methods Can Be Provably Faster than SGD after Finite Epochs 12 Jun 2020 · 0 repositories · arXiv:2006.07037
-
O(1) Communication for Distributed SGD through Two-Level Gradient Averaging 12 Jun 2020 · 0 repositories · arXiv:2006.07405
-
SGD with shuffling: optimal rates without component convexity and large epoch requirements 12 Jun 2020 · 0 repositories · arXiv:2006.06946
-
Stability of Stochastic Gradient Descent on Nonsmooth Convex Losses 12 Jun 2020 · 0 repositories · arXiv:2006.06914
-
Adaptive Gradient Methods Converge Faster with Over-Parameterization (but you should do a line-search) 11 Jun 2020 · 1 repository · arXiv:2006.06835Syntology 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
AdaS: Adaptive Scheduling of Stochastic Gradients 11 Jun 2020 · 2 repositories · arXiv:2006.06587
-
Borrowing From the Future: Addressing Double Sampling in Model-free Control 11 Jun 2020 · 0 repositories · arXiv:2006.06173
-
Multiplicative noise and heavy tails in stochastic optimization 11 Jun 2020 · 0 repositories · arXiv:2006.06293
-
Non-Convex SGD Learns Halfspaces with Adversarial Label Noise 11 Jun 2020 · 0 repositories · arXiv:2006.06742
-
STL-SGD: Speeding Up Local SGD with Stagewise Communication Period 11 Jun 2020 · 0 repositories · arXiv:2006.06377
-
Unifying Regularisation Methods for Continual Learning 11 Jun 2020 · 2 repositories · arXiv:2006.06357
-
Dynamical mean-field theory for stochastic gradient descent in Gaussian mixture classification 10 Jun 2020 · 0 repositories · arXiv:2006.06098
-
Random Reshuffling: Simple Analysis with Vast Improvements 10 Jun 2020 · 1 repository · arXiv:2006.05988Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 1 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Sketchy Empirical Natural Gradient Methods for Deep Learning 10 Jun 2020 · 1 repository · arXiv:2006.05924
-
Minibatch vs Local SGD for Heterogeneous Distributed Learning 8 Jun 2020 · 0 repositories · arXiv:2006.04735
-
Beyond Worst-Case Analysis in Stochastic Approximation: Moment Estimation Improves Instance Complexity 8 Jun 2020 · 0 repositories · arXiv:2006.04429
-
The Heavy-Tail Phenomenon in SGD 8 Jun 2020 · 1 repository · arXiv:2006.04740
-
The Strength of Nesterov's Extrapolation in the Individual Convergence of Nonsmooth Optimization 8 Jun 2020 · 0 repositories · arXiv:2006.04340
-
An Efficient Algorithm For Generalized Linear Bandit: Online Stochastic Gradient Descent and Thompson Sampling 7 Jun 2020 · 0 repositories · arXiv:2006.04012
-
Bayesian Neural Network via Stochastic Gradient Descent 4 Jun 2020 · 0 repositories · arXiv:2006.08453
-
Scaling Distributed Training with Adaptive Summation 4 Jun 2020 · 0 repositories · arXiv:2006.02924
-
Asymptotic Analysis of Conditioned Stochastic Gradient Descent 4 Jun 2020 · 1 repository · arXiv:2006.02745
-
Local SGD With a Communication Overhead Depending Only on the Number of Workers 3 Jun 2020 · 0 repositories · arXiv:2006.02582
-
On the Promise of the Stochastic Generalized Gauss-Newton Method for Training DNNs 3 Jun 2020 · 2 repositories · arXiv:2006.02409
-
ADAHESSIAN: An Adaptive Second Order Optimizer for Machine Learning 1 Jun 2020 · 4 repositories · arXiv:2006.00719Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
Augment Your Batch: Improving Generalization Through Instance Repetition 1 Jun 2020 · 0 repositories
-
Auto-Tuning Structured Light by Optical Stochastic Gradient Descent 1 Jun 2020 · 0 repositories
-
DaSGD: Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging 31 May 2020 · 0 repositories · arXiv:2006.00441
-
Inherent Noise in Gradient Based Methods 26 May 2020 · 0 repositories · arXiv:2005.12743
-
Microphone Array Based Surveillance Audio Classification 22 May 2020 · 0 repositories · arXiv:2005.11348
-
Accelerated Convergence for Counterfactual Learning to Rank 21 May 2020 · 1 repository · arXiv:2005.10615
-
rTop-k: A Statistical Estimation Approach to Distributed SGD 21 May 2020 · 0 repositories · arXiv:2005.10761
-
Stochastic Optimization with Heavy-Tailed Noise via Accelerated Gradient Clipping 21 May 2020 · 1 repository · arXiv:2005.10785Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Byzantine-Resilient SGD in High Dimensions on Heterogeneous Data 16 May 2020 · 0 repositories · arXiv:2005.07866
-
Learning the gravitational force law and other analytic functions 15 May 2020 · 0 repositories · arXiv:2005.07724
-
OD-SGD: One-step Delay Stochastic Gradient Descent for Distributed Training 14 May 2020 · 1 repository · arXiv:2005.06728
-
SQuARM-SGD: Communication-Efficient Momentum SGD for Decentralized Optimization 13 May 2020 · 0 repositories · arXiv:2005.07041
-
Convergence of Online Adaptive and Recurrent Optimization Algorithms 12 May 2020 · 0 repositories · arXiv:2005.05645
-
RSO: A Gradient Free Sampling Based Approach For Training Deep Neural Networks 12 May 2020 · 0 repositories · arXiv:2005.05955
-
Geoopt: Riemannian Optimization in PyTorch 6 May 2020 · 2 repositories · arXiv:2005.02819
-
Adaptive Learning of the Optimal Batch Size of SGD 3 May 2020 · 0 repositories · arXiv:2005.01097
-
Riemannian Stochastic Proximal Gradient Methods for Nonsmooth Optimization over the Stiefel Manifold 3 May 2020 · 0 repositories · arXiv:2005.01209
-
Gap-Aware Mitigation of Gradient Staleness 1 May 2020 · 0 repositories
-
Breaking (Global) Barriers in Parallel Stochastic Optimization with Wait-Avoiding Group Averaging 30 Apr 2020 · 0 repositories · arXiv:2005.00124
-
Learning Polynomials of Few Relevant Dimensions 28 Apr 2020 · 0 repositories · arXiv:2004.13748
-
The Impact of the Mini-batch Size on the Variance of Gradients in Stochastic Gradient Descent 27 Apr 2020 · 0 repositories · arXiv:2004.13146
-
Federated Learning with Only Positive Labels 21 Apr 2020 · 1 repository · arXiv:2004.10342
-
Stochastic gradient algorithms from ODE splitting perspective 19 Apr 2020 · 0 repositories · arXiv:2004.08981
-
On Tight Convergence Rates of Without-replacement SGD 18 Apr 2020 · 0 repositories · arXiv:2004.08657