Methods › General › Stochastic Optimization › SGD › Papers, page 12
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 12 of 21: papers 1,101 to 1,200 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Can Single-Shuffle SGD be Better than Reshuffling SGD and GD? 12 Mar 2021 · 0 repositories · arXiv:2103.07079
-
Streaming Linear System Identification with Reverse Experience Replay 10 Mar 2021 · 0 repositories · arXiv:2103.05896
-
Why flatness does and does not correlate with generalization for deep neural networks 10 Mar 2021 · 0 repositories · arXiv:2103.06219
-
Escaping Saddle Points with Stochastically Controlled Stochastic Gradient Methods 7 Mar 2021 · 0 repositories · arXiv:2103.04413
-
Second-order step-size tuning of SGD for non-convex optimization 5 Mar 2021 · 1 repository · arXiv:2103.03570
-
Better SGD using Second-order Momentum 4 Mar 2021 · 1 repository · arXiv:2103.03265
-
PAC-Bayes and Information Complexity 4 Mar 2021 · 0 repositories
-
A Biased Graph Neural Network Sampler with Near-Optimal Regret 1 Mar 2021 · 1 repository · arXiv:2103.01089
-
Non-Euclidean Differentially Private Stochastic Convex Optimization: Optimal Rates in Linear Time 1 Mar 2021 · 0 repositories · arXiv:2103.01278
-
Deep Neural Networks with ReLU-Sine-Exponential Activations Break Curse of Dimensionality in Approximation on Hölder Class 28 Feb 2021 · 0 repositories · arXiv:2103.00542
-
On the Utility of Gradient Compression in Distributed Training Systems 28 Feb 2021 · 1 repository · arXiv:2103.00543
-
Experiments with Rich Regime Training for Deep Learning 26 Feb 2021 · 0 repositories · arXiv:2102.13522
-
Noisy Truncated SGD: Optimization and Generalization 26 Feb 2021 · 0 repositories · arXiv:2103.00075
-
On the Generalization of Stochastic Gradient Descent with Momentum 26 Feb 2021 · 0 repositories · arXiv:2102.13653
-
Local Stochastic Gradient Descent Ascent: Convergence Analysis and Communication Efficiency 25 Feb 2021 · 0 repositories · arXiv:2102.13152
-
Loss Surface Simplexes for Mode Connecting Volumes and Fast Ensembling 25 Feb 2021 · 1 repository · arXiv:2102.13042
-
Machine Unlearning via Algorithmic Stability 25 Feb 2021 · 0 repositories · arXiv:2102.13179
-
Mixed Variable Bayesian Optimization with Frequency Modulated Kernels 25 Feb 2021 · 0 repositories · arXiv:2102.12792
-
On the Validity of Modeling SGD with Stochastic Differential Equations (SDEs) 24 Feb 2021 · 1 repository · arXiv:2102.12470
-
Online Stochastic Gradient Descent Learns Linear Dynamical Systems from A Single Trajectory 23 Feb 2021 · 0 repositories · arXiv:2102.11822
-
Convergence Rates of Stochastic Gradient Descent under Infinite Noise Variance 20 Feb 2021 · 0 repositories · arXiv:2102.10346
-
Permutation-Based SGD: Is Random Optimal? 19 Feb 2021 · 1 repository · arXiv:2102.09718Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 4 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Personalized Federated Learning: A Unified Framework and Universal Optimization Techniques 19 Feb 2021 · 1 repository · arXiv:2102.09743
-
SVRG Meets AdaGrad: Painless Variance Reduction 18 Feb 2021 · 0 repositories · arXiv:2102.09645
-
Proactive DP: A Multple Target Optimization Framework for DP-SGD 17 Feb 2021 · 0 repositories · arXiv:2102.09030
-
Genetically Optimized Prediction of Remaining Useful Life 17 Feb 2021 · 0 repositories · arXiv:2102.08845
-
Convergence of stochastic gradient descent schemes for Lojasiewicz-landscapes 16 Feb 2021 · 0 repositories · arXiv:2102.09385
-
Differential Privacy and Byzantine Resilience in SGD: Do They Add Up? 16 Feb 2021 · 1 repository · arXiv:2102.08166
-
GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training 16 Feb 2021 · 2 repositories · arXiv:2102.08098Syntology official (archive's flag): 6 ran · 7 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 3 where Syntology's instrument failed) · 6 unverified (of 13 harvested samples) · 12 pointer-only (licence)
-
IntSGD: Adaptive Floatless Compression of Stochastic Gradients 16 Feb 2021 · 1 repository · arXiv:2102.08374Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Does the Adam Optimizer Exacerbate Catastrophic Forgetting? 15 Feb 2021 · 1 repository · arXiv:2102.07686
-
The Role of Momentum Parameters in the Optimal Convergence of Adaptive Polyak's Heavy-ball Methods 15 Feb 2021 · 0 repositories · arXiv:2102.07314
-
Learning by Turning: Neural Architecture Aware Optimisation 14 Feb 2021 · 2 repositories · arXiv:2102.07227Syntology official (archive's flag): 2 ran · 5 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
Asymmetric Heavy Tails and Implicit Bias in Gaussian Noise Injections 13 Feb 2021 · 1 repository · arXiv:2102.07006
-
Bayesian Neural Network Priors Revisited 12 Feb 2021 · 1 repository · arXiv:2102.06571
-
Proximal and Federated Random Reshuffling 12 Feb 2021 · 1 repository · arXiv:2102.06704Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Stability and Convergence of Stochastic Gradient Clipping: Beyond Lipschitz Continuity and Smoothness 12 Feb 2021 · 0 repositories · arXiv:2102.06489
-
Strength of Minibatch Noise in SGD 10 Feb 2021 · 0 repositories · arXiv:2102.05375
-
On PyTorch Implementation of Density Estimators for von Mises-Fisher and Its Mixture 10 Feb 2021 · 1 repository · arXiv:2102.05340
-
Federated Learning with Local Differential Privacy: Trade-offs between Privacy, Utility, and Communication 9 Feb 2021 · 0 repositories · arXiv:2102.04737
-
Coordinating Momenta for Cross-silo Federated Learning 8 Feb 2021 · 0 repositories · arXiv:2102.03970
-
Eliminating Sharp Minima from SGD with Truncated Heavy-tailed Noise 8 Feb 2021 · 0 repositories · arXiv:2102.04297
-
SGD in the Large: Average-case Analysis, Asymptotics, and Stepsize Criticality 8 Feb 2021 · 0 repositories · arXiv:2102.04396
-
Dimension Free Generalization Bounds for Non Linear Metric Learning 7 Feb 2021 · 0 repositories · arXiv:2102.03802
-
Weight Rescaling: Effective and Robust Regularization for Deep Neural Networks with Batch Normalization 6 Feb 2021 · 0 repositories · arXiv:2102.03497
-
Bias-Variance Reduced Local SGD for Less Heterogeneous Federated Learning 5 Feb 2021 · 0 repositories · arXiv:2102.03198
-
Fast and Memory Efficient Differentially Private-SGD via JL Projections 5 Feb 2021 · 0 repositories · arXiv:2102.03013
-
Last iterate convergence of SGD for Least-Squares in the Interpolation regime 5 Feb 2021 · 0 repositories · arXiv:2102.03183
-
Generalization Bounds for Noisy Iterative Algorithms Using Properties of Additive Noise Channels 5 Feb 2021 · 0 repositories · arXiv:2102.02976
-
1-bit Adam: Communication Efficient Large-Scale Training with Adam's Convergence Speed 4 Feb 2021 · 2 repositories · arXiv:2102.02888
-
Stability and Generalization of the Decentralized Stochastic Gradient Descent 2 Feb 2021 · 0 repositories · arXiv:2102.01302
-
Information-Theoretic Generalization Bounds for Stochastic Gradient Descent 1 Feb 2021 · 0 repositories · arXiv:2102.00931
-
SGD Generalizes Better Than GD (And Regularization Doesn't Help) 1 Feb 2021 · 0 repositories · arXiv:2102.01117
-
On the Origin of Implicit Regularization in Stochastic Gradient Descent 28 Jan 2021 · 0 repositories · arXiv:2101.12176
-
Adaptivity without Compromise: A Momentumized, Adaptive, Dual Averaged Gradient Method for Stochastic Optimization 26 Jan 2021 · 5 repositories · arXiv:2101.11075
-
Differentially Private SGD with Non-Smooth Losses 22 Jan 2021 · 0 repositories · arXiv:2101.08925
-
Clairvoyant Prefetching for Distributed Machine Learning I/O 21 Jan 2021 · 0 repositories · arXiv:2101.08734
-
Sum-Rate-Distortion Function for Indirect Multiterminal Source Coding in Federated Learning 21 Jan 2021 · 0 repositories · arXiv:2101.08696
-
Guided parallelized stochastic gradient descent for delay compensation 17 Jan 2021 · 0 repositories · arXiv:2101.07259
-
Linguistically-Enriched and Context-Aware Zero-shot Slot Filling 16 Jan 2021 · 0 repositories · arXiv:2101.06514
-
Phases of learning dynamics in artificial neural networks: with or without mislabeled data 16 Jan 2021 · 0 repositories · arXiv:2101.06509
-
BN-invariant sharpness regularizes the training model to better generalization 8 Jan 2021 · 0 repositories · arXiv:2101.02944
-
Towards Understanding Learning in Neural Networks with Linear Teachers 7 Jan 2021 · 0 repositories · arXiv:2101.02533
-
Can Transfer Neuroevolution Tractably Solve Your Differential Equations? 6 Jan 2021 · 0 repositories · arXiv:2101.01998
-
Federated Learning over Noisy Channels: Convergence Analysis and Design Examples 6 Jan 2021 · 0 repositories · arXiv:2101.02198
-
Provable Generalization of SGD-trained Neural Networks of Any Width in the Presence of Adversarial Label Noise 4 Jan 2021 · 1 repository · arXiv:2101.01152
-
A Chaos Theory Approach to Understand Neural Network Optimization 1 Jan 2021 · 0 repositories
-
A New Variant of Stochastic Heavy ball Optimization Method for Deep Learning 1 Jan 2021 · 0 repositories
-
Accelerating DNN Training through Selective Localized Learning 1 Jan 2021 · 0 repositories
-
Adaptive Dataset Sampling by Deep Policy Gradient 1 Jan 2021 · 0 repositories
-
Adaptive Gradient Methods Can Be Provably Faster than SGD with Random Shuffling 1 Jan 2021 · 0 repositories
-
Apollo: An Adaptive Parameter-wised Diagonal Quasi-Newton Method for Nonconvex Stochastic Optimization 1 Jan 2021 · 0 repositories
-
Compressing gradients in distributed SGD by exploiting their temporal correlation 1 Jan 2021 · 0 repositories
-
Constructing Multiple High-Quality Deep Neural Networks: A TRUST-TECH Based Approach 1 Jan 2021 · 0 repositories
-
DQSGD: DYNAMIC QUANTIZED STOCHASTIC GRADIENT DESCENT FOR COMMUNICATION-EFFICIENT DISTRIBUTED LEARNING 1 Jan 2021 · 0 repositories
-
Factor Normalization for Deep Neural Network Models 1 Jan 2021 · 1 repository
-
Fast convergence of stochastic subgradient method under interpolation 1 Jan 2021 · 0 repositories
-
Flatness is a Flase Friend 1 Jan 2021 · 0 repositories
-
How Benign is Benign Overfitting ? 1 Jan 2021 · 0 repositories
-
Implicit Regularization Effects of Unbiased Random Label Noises with SGD 1 Jan 2021 · 0 repositories
-
Implicit Regularization of SGD via Thermophoresis 1 Jan 2021 · 0 repositories
-
Local SGD Meets Asynchrony 1 Jan 2021 · 0 repositories
-
Multi-Level Local SGD: Distributed SGD for Heterogeneous Hierarchical Networks 1 Jan 2021 · 0 repositories
-
Noise against noise: stochastic label noise helps combat inherent label noise 1 Jan 2021 · 0 repositories
-
On the Inductive Bias of a CNN for Distributions with Orthogonal Patterns 1 Jan 2021 · 0 repositories
-
Optimizing Quantized Neural Networks with Natural Gradient 1 Jan 2021 · 0 repositories
-
Revisiting the Stability of Stochastic Gradient Descent: A Tightness Analysis 1 Jan 2021 · 0 repositories
-
Stochastic Optimization with Non-stationary Noise: The Power of Moment Estimation 1 Jan 2021 · 0 repositories
-
Stochastic Proximal Point Algorithm for Large-scale Nonconvex Optimization: Convergence, Implementation, and Application to Neural Networks 1 Jan 2021 · 0 repositories
-
Symmetry, Conservation Laws, and Learning Dynamics in Neural Networks 1 Jan 2021 · 0 repositories
-
The Impact of the Mini-batch Size on the Dynamics of SGD: Variance and Beyond 1 Jan 2021 · 0 repositories
-
The simpler the better: vanilla sgd revisited 1 Jan 2021 · 0 repositories
-
Why Does Decentralized Training Outperform Synchronous Training In The Large Batch Setting? 1 Jan 2021 · 0 repositories
-
SGD Distributional Dynamics of Three Layer Neural Networks 30 Dec 2020 · 0 repositories · arXiv:2012.15036
-
Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts Generalization 28 Dec 2020 · 0 repositories · arXiv:2012.14193
-
AsymptoticNG: A regularized natural gradient optimization algorithm with look-ahead strategy 24 Dec 2020 · 0 repositories · arXiv:2012.13077
-
Training Convolutional Neural Networks With Hebbian Principal Component Analysis 22 Dec 2020 · 1 repository · arXiv:2012.12229
-
Optimizing Deep Neural Networks through Neuroevolution with Stochastic Gradient Descent 21 Dec 2020 · 0 repositories · arXiv:2012.11184
-
A hybrid MGA-MSGD ANN training approach for approximate solution of linear elliptic PDEs 18 Dec 2020 · 0 repositories · arXiv:2012.11517
-
FedADC: Accelerated Federated Learning with Drift Control 16 Dec 2020 · 0 repositories · arXiv:2012.09102