Methods › General › Stochastic Optimization › SGD › Papers, page 10
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 10 of 21: papers 901 to 1,000 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Dynamic Differential-Privacy Preserving SGD 30 Oct 2021 · 0 repositories · arXiv:2111.00173
-
Does Momentum Help? A Sample Complexity Analysis 29 Oct 2021 · 0 repositories · arXiv:2110.15547
-
Eigencurve: Optimal Learning Rate Schedule for SGD on Quadratic Objectives with Skewed Hessian Spectrums 27 Oct 2021 · 1 repository · arXiv:2110.14109
-
Multilayer Lookahead: a Nested Version of Lookahead 27 Oct 2021 · 0 repositories · arXiv:2110.14254
-
Exponential Graph is Provably Efficient for Decentralized Deep Training 26 Oct 2021 · 2 repositories · arXiv:2110.13363Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Asynchronous Decentralized Distributed Training of Acoustic Models 21 Oct 2021 · 0 repositories · arXiv:2110.11199
-
Towards Noise-adaptive, Problem-adaptive (Accelerated) Stochastic Gradient Descent 21 Oct 2021 · 0 repositories · arXiv:2110.11442
-
Minibatch vs Local SGD with Shuffling: Tight Convergence Bounds and Beyond 20 Oct 2021 · 0 repositories · arXiv:2110.10342
-
Training Deep Neural Networks with Adaptive Momentum Inspired by the Quadratic Optimization 18 Oct 2021 · 1 repository · arXiv:2110.09057Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 1 honoured, 3 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 17 harvested samples) · 5 pointer-only (licence)
-
Improving Compositional Generalization with Self-Training for Data-to-Text Generation 16 Oct 2021 · 1 repository · arXiv:2110.08467
-
Towards Better Plasticity-Stability Trade-off in Incremental Learning: A Simple Linear Connector 15 Oct 2021 · 0 repositories · arXiv:2110.07905
-
Trade-offs of Local SGD at Scale: An Empirical Study 15 Oct 2021 · 0 repositories · arXiv:2110.08133
-
Adaptive Elastic Training for Sparse Deep Learning on Heterogeneous Multi-GPU Servers 13 Oct 2021 · 1 repository · arXiv:2110.07029
-
On the Double Descent of Random Features Models Trained with SGD 13 Oct 2021 · 0 repositories · arXiv:2110.06910
-
SGD-X: A Benchmark for Robust Generalization in Schema-Guided Dialogue Systems 13 Oct 2021 · 1 repository · arXiv:2110.06800
-
What Happens after SGD Reaches Zero Loss? --A Mathematical Framework 13 Oct 2021 · 0 repositories · arXiv:2110.06914
-
Last Iterate Risk Bounds of SGD with Decaying Stepsize for Overparameterized Linear Regression 12 Oct 2021 · 0 repositories · arXiv:2110.06198
-
Not all noise is accounted equally: How differentially private learning benefits from large sampling rates 12 Oct 2021 · 1 repository · arXiv:2110.06255
-
The Role of Permutation Invariance in Linear Mode Connectivity of Neural Networks 12 Oct 2021 · 1 repository · arXiv:2110.06296
-
Momentum Centering and Asynchronous Update for Adaptive Gradient Methods 11 Oct 2021 · 2 repositories · arXiv:2110.05454
-
Frequency-aware SGD for Efficient Embedding Learning with Provable Benefits 10 Oct 2021 · 0 repositories · arXiv:2110.04844
-
Combining Differential Privacy and Byzantine Resilience in Distributed SGD 8 Oct 2021 · 0 repositories · arXiv:2110.03991
-
Does Momentum Change the Implicit Regularization on Separable Data? 8 Oct 2021 · 0 repositories · arXiv:2110.03891
-
On the Generalization of Models Trained with SGD: Information-Theoretic Bounds and Implications 7 Oct 2021 · 0 repositories · arXiv:2110.03128
-
Spectral Bias in Practice: The Role of Function Frequency in Generalization 6 Oct 2021 · 0 repositories · arXiv:2110.02424
-
Rapid training of deep neural networks without skip connections or normalization layers using Deep Kernel Shaping 5 Oct 2021 · 2 repositories · arXiv:2110.01765
-
S2 Reducer: High-Performance Sparse Communication to Accelerate Distributed Deep Learning 5 Oct 2021 · 0 repositories · arXiv:2110.02140
-
Effectiveness of Optimization Algorithms in Deep Image Classification 4 Oct 2021 · 1 repository · arXiv:2110.01598
-
Global Convergence and Stability of Stochastic Gradient Descent 4 Oct 2021 · 0 repositories · arXiv:2110.01663
-
Layer-wise and Dimension-wise Locally Adaptive Federated Learning 1 Oct 2021 · 0 repositories · arXiv:2110.00532
-
A Class of Short-term Recurrence Anderson Mixing Methods and Their Applications 29 Sep 2021 · 0 repositories
-
A General Analysis of Example-Selection for Stochastic Gradient Descent 29 Sep 2021 · 0 repositories
-
A Novel Convergence Analysis for the Stochastic Proximal Point Algorithm 29 Sep 2021 · 0 repositories
-
A Practical PAC-Bayes Generalisation Bound for Deep Learning 29 Sep 2021 · 0 repositories
-
Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and Momentum 29 Sep 2021 · 0 repositories
-
Boosting the Confidence of Near-Tight Generalization Bounds for Uniformly Stable Randomized Algorithms 29 Sep 2021 · 0 repositories
-
Determining the Ethno-nationality of Writers Using Written English Text 29 Sep 2021 · 0 repositories
-
Directional Bias Helps Stochastic Gradient Descent to Generalize in Nonparametric Model 29 Sep 2021 · 0 repositories
-
Efficient Second-Order Optimization for Deep Learning with Kernel Machines 29 Sep 2021 · 0 repositories
-
How to Improve Sample Complexity of SGD over Highly Dependent Data? 29 Sep 2021 · 0 repositories
-
Hybrid Local SGD for Federated Learning with Heterogeneous Communications 29 Sep 2021 · 0 repositories
-
Image Functions In Neural Networks: A Perspective On Generalization 29 Sep 2021 · 0 repositories
-
Learning Pruning-Friendly Networks via Frank-Wolfe: One-Shot, Any-Sparsity, And No Retraining 29 Sep 2021 · 1 repository
-
Learning Rate Grafting: Transferability of Optimizer Tuning 29 Sep 2021 · 0 repositories
-
Linear Backpropagation Leads to Faster Convergence 29 Sep 2021 · 0 repositories
-
Logarithmic landscape and power-law escape rate of SGD 29 Sep 2021 · 0 repositories
-
On the Convergence of Nonconvex Continual Learning with Adaptive Learning Rate 29 Sep 2021 · 0 repositories
-
Optimization and Adaptive Generalization of Three layer Neural Networks 29 Sep 2021 · 0 repositories
-
Orthogonalising gradients to speedup neural network optimisation 29 Sep 2021 · 0 repositories
-
Rethinking the limiting dynamics of SGD: modified loss, phase space oscillations, and anomalous diffusion 29 Sep 2021 · 0 repositories
-
SGD Can Converge to Local Maxima 29 Sep 2021 · 0 repositories
-
SLIM-QN: A Stochastic, Light, Momentumized Quasi-Newton Optimizer for Deep Neural Networks 29 Sep 2021 · 0 repositories
-
SpSC: A Fast and Provable Algorithm for Sampling-Based GNN Training 29 Sep 2021 · 0 repositories
-
Squeezing SGD Parallelization Performance in Distributed Training Using Delayed Averaging 29 Sep 2021 · 0 repositories
-
Stochastic Training is Not Necessary for Generalization 29 Sep 2021 · 1 repository · arXiv:2109.14119
-
Two Regimes of Generalization for Non-Linear Metric Learning 29 Sep 2021 · 0 repositories
-
IGLU: Efficient GCN Training via Lazy Updates 28 Sep 2021 · 1 repository · arXiv:2109.13995Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Accelerated PDEs for Construction and Theoretical Analysis of an SGD Extension 27 Sep 2021 · 0 repositories
-
Unrolling SGD: Understanding Factors Influencing Machine Unlearning 27 Sep 2021 · 1 repository · arXiv:2109.13398Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
AdaInject: Injection Based Adaptive Gradient Descent Optimizers for Convolutional Neural Networks 26 Sep 2021 · 1 repository · arXiv:2109.12504
-
MixNN: Protection of Federated Learning Against Inference Attacks by Mixing Neural Network Layers 26 Sep 2021 · 0 repositories · arXiv:2109.12550
-
Unbiased Single-scale and Multi-scale Quantizers for Distributed Optimization 26 Sep 2021 · 1 repository · arXiv:2109.12497
-
NanoBatch Privacy: Enabling fast Differentially Private learning on the IPU 24 Sep 2021 · 0 repositories · arXiv:2109.12191
-
Neural network relief: a pruning algorithm based on neural activity 22 Sep 2021 · 1 repository · arXiv:2109.10795
-
Self-learn to Explain Siamese Networks Robustly 15 Sep 2021 · 0 repositories · arXiv:2109.07371
-
Toward Communication Efficient Adaptive Gradient Method 10 Sep 2021 · 0 repositories · arXiv:2109.05109
-
Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time Analysis 8 Sep 2021 · 0 repositories · arXiv:2109.03699
-
COCO Denoiser: Using Co-Coercivity for Variance Reduction in Stochastic Convex Optimization 7 Sep 2021 · 1 repository · arXiv:2109.03207
-
Revisiting Recursive Least Squares for Training Deep Neural Networks 7 Sep 2021 · 0 repositories · arXiv:2109.03220
-
Stochastic Subgradient Descent on a Generic Definable Function Converges to a Minimizer 6 Sep 2021 · 0 repositories · arXiv:2109.02455
-
Statistical Estimation and Inference via Local SGD in Federated Learning 3 Sep 2021 · 0 repositories · arXiv:2109.01326
-
The Minimax Complexity of Distributed Optimization 1 Sep 2021 · 0 repositories · arXiv:2109.00534
-
Using a one dimensional parabolic model of the full-batch loss to estimate learning rates during training 31 Aug 2021 · 1 repository · arXiv:2108.13880
-
Byzantine Fault-Tolerance in Federated Local SGD under 2f-Redundancy 26 Aug 2021 · 0 repositories · arXiv:2108.11769
-
Shift-Curvature, SGD, and Generalization 21 Aug 2021 · 0 repositories · arXiv:2108.09507
-
EDEN: Communication-Efficient and Robust Distributed Mean Estimation for Federated Learning 19 Aug 2021 · 1 repository · arXiv:2108.08842Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Compressing gradients by exploiting temporal correlation in momentum-SGD 17 Aug 2021 · 0 repositories · arXiv:2108.07827
-
FedPAGE: A Fast Local Stochastic Gradient Method for Communication-Efficient Federated Learning 10 Aug 2021 · 0 repositories · arXiv:2108.04755
-
The Benefits of Implicit Regularization from SGD in Least Squares Problems 10 Aug 2021 · 0 repositories · arXiv:2108.04552
-
Efficient Hyperparameter Optimization for Differentially Private Deep Learning 9 Aug 2021 · 1 repository · arXiv:2108.03888
-
On the Hyperparameters in Stochastic Gradient Descent with Momentum 9 Aug 2021 · 0 repositories · arXiv:2108.03947
-
On the Power of Differentiable Learning versus PAC and SQ Learning 9 Aug 2021 · 0 repositories · arXiv:2108.04190
-
Unified Regularity Measures for Sample-wise Learning and Generalization 9 Aug 2021 · 0 repositories · arXiv:2108.03913
-
Stochastic Subgradient Descent Escapes Active Strict Saddles on Weakly Convex Functions 4 Aug 2021 · 0 repositories · arXiv:2108.02072
-
Large-Scale Differentially Private BERT 3 Aug 2021 · 0 repositories · arXiv:2108.01624
-
Rethinking gradient sparsification as total error minimization 2 Aug 2021 · 0 repositories · arXiv:2108.00951
-
How much pre-training is enough to discover a good subnetwork? 31 Jul 2021 · 0 repositories · arXiv:2108.00259
-
DQ-SGD: Dynamic Quantization in SGD for Communication-Efficient Distributed Learning 30 Jul 2021 · 0 repositories · arXiv:2107.14575
-
Manipulating Identical Filter Redundancy for Efficient Pruning on Deep and Complicated CNN 30 Jul 2021 · 2 repositories · arXiv:2107.14444
-
Decentralized Federated Learning: Balancing Communication and Computing Costs 26 Jul 2021 · 1 repository · arXiv:2107.12048
-
SGD with a Constant Large Learning Rate Can Converge to Local Maxima 25 Jul 2021 · 0 repositories · arXiv:2107.11774
-
A general sample complexity analysis of vanilla policy gradient 23 Jul 2021 · 0 repositories · arXiv:2107.11433
-
Local SGD Optimizes Overparameterized Neural Networks in Polynomial Time 22 Jul 2021 · 0 repositories · arXiv:2107.10868
-
Distribution of Classification Margins: Are All Data Equal? 21 Jul 2021 · 0 repositories · arXiv:2107.10199
-
Improved Learning Rates for Stochastic Optimization: Two Theoretical Viewpoints 19 Jul 2021 · 0 repositories · arXiv:2107.08686
-
Non-asymptotic estimates for TUSLA algorithm for non-convex learning with applications to neural networks with ReLU activation function 19 Jul 2021 · 1 repository · arXiv:2107.08649
-
The Limiting Dynamics of SGD: Modified Loss, Phase Space Oscillations, and Anomalous Diffusion 19 Jul 2021 · 1 repository · arXiv:2107.09133
-
A New Adaptive Gradient Method with Gradient Decomposition 18 Jul 2021 · 0 repositories · arXiv:2107.08377
-
Accelerating Distributed K-FAC with Smart Parallelism of Computing and Communication Tasks 14 Jul 2021 · 0 repositories · arXiv:2107.06533
-
SGD: The Role of Implicit Regularization, Batch-size and Multiple-epochs 11 Jul 2021 · 0 repositories · arXiv:2107.05074