Methods › General › Stochastic Optimization › SGD › Papers, page 8
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 8 of 21: papers 701 to 800 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Is Stochastic Gradient Descent Near Optimal? 18 Sep 2022 · 0 repositories · arXiv:2209.08627
-
Renyi Differential Privacy of Propose-Test-Release and Applications to Private and Robust Machine Learning 16 Sep 2022 · 0 repositories · arXiv:2209.07716
-
Efficiency Ordering of Stochastic Gradient Descent 15 Sep 2022 · 0 repositories · arXiv:2209.07446
-
Random initialisations performing above chance and how to find them 15 Sep 2022 · 1 repository · arXiv:2209.07509
-
Optimization without Backpropagation 13 Sep 2022 · 1 repository · arXiv:2209.06302Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 8 harvested samples)
-
Personalized Federated Learning with Communication Compression 12 Sep 2022 · 0 repositories · arXiv:2209.05148
-
Differentially Private Stochastic Gradient Descent with Low-Noise 9 Sep 2022 · 0 repositories · arXiv:2209.04188
-
Revisiting Outer Optimization in Adversarial Training 2 Sep 2022 · 0 repositories · arXiv:2209.01199
-
Meta Objective Guided Disambiguation for Partial Label Learning 26 Aug 2022 · 0 repositories · arXiv:2208.12459
-
A simplified convergence theory for Byzantine resilient stochastic gradient descent 25 Aug 2022 · 0 repositories · arXiv:2208.11879
-
Accelerating SGD for Highly Ill-Conditioned Huge-Scale Online Matrix Completion 24 Aug 2022 · 1 repository · arXiv:2208.11246
-
Robustness to Unbounded Smoothness of Generalized SignSGD 23 Aug 2022 · 0 repositories · arXiv:2208.11195
-
Provable Adaptivity of Adam under Non-uniform Smoothness 21 Aug 2022 · 0 repositories · arXiv:2208.09900
-
Learning with Local Gradients at the Edge 17 Aug 2022 · 0 repositories · arXiv:2208.08503
-
Training Overparametrized Neural Networks in Sublinear Time 9 Aug 2022 · 0 repositories · arXiv:2208.04508
-
Adaptive Stochastic Gradient Descent for Fast and Communication-Efficient Distributed Learning 4 Aug 2022 · 0 repositories · arXiv:2208.03134
-
Feature selection with gradient descent on two-layer networks in low-rotation regimes 4 Aug 2022 · 0 repositories · arXiv:2208.02789
-
Formal guarantees for heuristic optimization algorithms used in machine learning 31 Jul 2022 · 0 repositories · arXiv:2208.00502
-
DRSOM: A Dimension Reduced Second-Order Method 30 Jul 2022 · 3 repositories · arXiv:2208.00208
-
CrAM: A Compression-Aware Minimizer 28 Jul 2022 · 1 repository · arXiv:2207.14200Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified; the one sample that ran constructed an object rather than computing a result (of 2 harvested samples)
-
DADAO: Decoupled Accelerated Decentralized Asynchronous Optimization 26 Jul 2022 · 1 repository · arXiv:2208.00779
-
Adaptive Step-Size Methods for Compressed SGD 20 Jul 2022 · 0 repositories · arXiv:2207.10046
-
Moment Centralization based Gradient Descent Optimizers for Convolutional Neural Networks 19 Jul 2022 · 1 repository · arXiv:2207.09066
-
Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit 18 Jul 2022 · 0 repositories · arXiv:2207.08799
-
Communication-efficient Distributed Learning for Large Batch Optimization 17 Jul 2022 · 1 repository
-
SP2: A Second Order Stochastic Polyak Method 17 Jul 2022 · 0 repositories · arXiv:2207.08171
-
Adaptive Sketches for Robust Regression with Importance Sampling 16 Jul 2022 · 0 repositories · arXiv:2207.07822
-
MixTailor: Mixed Gradient Aggregation for Robust Learning Against Tailored Attacks 16 Jul 2022 · 0 repositories · arXiv:2207.07941
-
TCT: Convexifying Federated Learning using Bootstrapped Neural Tangent Kernels 13 Jul 2022 · 1 repository · arXiv:2207.06343Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Towards understanding how momentum improves generalization in deep learning 13 Jul 2022 · 0 repositories · arXiv:2207.05931
-
Stochastic Gradient Descent and Anomaly of Variance-flatness Relation in Artificial Neural Networks 11 Jul 2022 · 0 repositories · arXiv:2207.04932
-
On uniform-in-time diffusion approximation for stochastic gradient descent 11 Jul 2022 · 0 repositories · arXiv:2207.04922
-
Challenging Common Assumptions about Catastrophic Forgetting 10 Jul 2022 · 0 repositories · arXiv:2207.04543
-
Improved Binary Forward Exploration: Learning Rate Scheduling Method for Stochastic Optimization 9 Jul 2022 · 0 repositories · arXiv:2207.04198
-
The Poisson binomial mechanism for secure and private federated learning 9 Jul 2022 · 0 repositories · arXiv:2207.09916
-
A Deep Model for Partial Multi-Label Image Classification with Curriculum Based Disambiguation 6 Jul 2022 · 0 repositories · arXiv:2207.02410
-
BFE and AdaBFE: A New Approach in Learning Rate Automation for Stochastic Optimization 6 Jul 2022 · 0 repositories · arXiv:2207.02763
-
The alignment property of SGD noise and how it helps select flat minima: A stability analysis 6 Jul 2022 · 0 repositories · arXiv:2207.02628
-
ST-CoNAL: Consistency-Based Acquisition Criterion Using Temporal Self-Ensemble for Active Learning 5 Jul 2022 · 0 repositories · arXiv:2207.02182
-
A Theoretical Analysis of the Learning Dynamics under Class Imbalance 1 Jul 2022 · 1 repository · arXiv:2207.00391
-
Studying Generalization Through Data Averaging 28 Jun 2022 · 0 repositories · arXiv:2206.13669
-
Making Look-Ahead Active Learning Strategies Feasible with Neural Tangent Kernels 25 Jun 2022 · 0 repositories · arXiv:2206.12569
-
Statistical inference with implicit SGD: proximal Robbins-Monro vs. Polyak-Ruppert 25 Jun 2022 · 0 repositories · arXiv:2206.12663
-
A view of mini-batch SGD via generating functions: conditions of convergence, phase transitions, benefit from negative momenta 22 Jun 2022 · 1 repository · arXiv:2206.11124Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
On the Maximum Hessian Eigenvalue and Generalization 21 Jun 2022 · 0 repositories · arXiv:2206.10654
-
Low-Precision Stochastic Gradient Langevin Dynamics 20 Jun 2022 · 1 repository · arXiv:2206.09909Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Provable Generalization of Overparameterized Meta-learning Trained with SGD 18 Jun 2022 · 0 repositories · arXiv:2206.09136
-
A Closer Look at Smoothness in Domain Adversarial Training 16 Jun 2022 · 1 repository · arXiv:2206.08213Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Sharper Convergence Guarantees for Asynchronous SGD for Distributed and Federated Learning 16 Jun 2022 · 0 repositories · arXiv:2206.08307
-
Asynchronous SGD Beats Minibatch SGD Under Arbitrary Delays 15 Jun 2022 · 1 repository · arXiv:2206.07638
-
Implicit Regularization or Implicit Conditioning? Exact Risk Trajectories of SGD in High Dimensions 15 Jun 2022 · 0 repositories · arXiv:2206.07252
-
Markov Chain Score Ascent: A Unifying Framework of Variational Inference with Markovian Gradients 13 Jun 2022 · 1 repository · arXiv:2206.06295
-
SGD and Weight Decay Secretly Minimize the Rank of Your Neural Network 12 Jun 2022 · 0 repositories · arXiv:2206.05794
-
Stochastic Gradient Descent without Full Data Shuffle 12 Jun 2022 · 1 repository · arXiv:2206.05830
-
Bayesian Estimation of Differential Privacy 10 Jun 2022 · 1 repository · arXiv:2206.05199Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Learning Non-Vacuous Generalization Bounds from Optimization 9 Jun 2022 · 1 repository · arXiv:2206.04359
-
High-dimensional limit theorems for SGD: Effective dynamics and critical scaling 8 Jun 2022 · 0 repositories · arXiv:2206.04030
-
Click prediction boosting via Bayesian hyperparameter optimization based ensemble learning pipelines 7 Jun 2022 · 0 repositories · arXiv:2206.03592
-
Generalization Error Bounds for Deep Neural Networks Trained by SGD 7 Jun 2022 · 0 repositories · arXiv:2206.03299
-
Integrating Random Effects in Deep Neural Networks 7 Jun 2022 · 1 repository · arXiv:2206.03314Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Early Stage Convergence and Global Convergence of Training Mildly Parameterized Neural Networks 5 Jun 2022 · 1 repository · arXiv:2206.02139
-
Sharper Rates and Flexible Framework for Nonconvex SGD with Client and Data Sampling 5 Jun 2022 · 1 repository · arXiv:2206.02275
-
A PDE-based Explanation of Extreme Numerical Sensitivities and Edge of Stability in Training Neural Networks 4 Jun 2022 · 1 repository · arXiv:2206.02001
-
Algorithmic Stability of Heavy-Tailed Stochastic Gradient Descent on Least Squares 2 Jun 2022 · 0 repositories · arXiv:2206.01274
-
Stochastic gradient descent introduces an effective landscape-dependent regularization favoring flat solutions 2 Jun 2022 · 0 repositories · arXiv:2206.01246
-
Trajectory of Mini-Batch Momentum: Batch Size Saturation and Convergence in High Dimensions 2 Jun 2022 · 0 repositories · arXiv:2206.01029
-
A Theoretical Framework for Inference Learning 1 Jun 2022 · 1 repository · arXiv:2206.00164Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Computing the Variance of Shuffling Stochastic Gradient Algorithms via Power Spectral Density Analysis 1 Jun 2022 · 1 repository · arXiv:2206.00632
-
Optimization with Access to Auxiliary Information 1 Jun 2022 · 1 repository · arXiv:2206.00395
-
Metrizing Fairness 30 May 2022 · 1 repository · arXiv:2205.15049
-
Special Properties of Gradient Descent with Large Learning Rates 30 May 2022 · 0 repositories · arXiv:2205.15142
-
Re-parameterizing Your Optimizers rather than Architectures 30 May 2022 · 1 repository · arXiv:2205.15242Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Generalization Bounds for Gradient Methods via Discrete and Continuous Prior 27 May 2022 · 0 repositories · arXiv:2205.13799
-
Privacy of Noisy Stochastic Gradient Descent: More Iterations without More Privacy Loss 27 May 2022 · 0 repositories · arXiv:2205.13710
-
FedBR: Improving Federated Learning on Heterogeneous Data via Local Learning Bias Reduction 26 May 2022 · 1 repository · arXiv:2205.13462Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 4 where Syntology's instrument failed) · 5 unverified (of 15 harvested samples) · 2 pointer-only (licence)
-
Trainable Weight Averaging: A General Approach for Subspace Training 26 May 2022 · 1 repository · arXiv:2205.13104Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 5 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Mirror Descent Maximizes Generalized Margin and Can Be Implemented Efficiently 25 May 2022 · 0 repositories · arXiv:2205.12808
-
Byzantine Machine Learning Made Easy by Resilient Averaging of Momentums 24 May 2022 · 0 repositories · arXiv:2205.12173
-
Do Deep Learning Models and News Headlines Outperform Conventional Prediction Techniques on Forex Data? 22 May 2022 · 0 repositories · arXiv:2205.10743
-
GraB: Finding Provably Better Data Permutations than Random Reshuffling 22 May 2022 · 3 repositories · arXiv:2205.10733Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
On the SDEs and Scaling Rules for Adaptive Gradient Algorithms 20 May 2022 · 1 repository · arXiv:2205.10287
-
Homogenization of SGD in high-dimensions: Exact dynamics and generalization properties 14 May 2022 · 0 repositories · arXiv:2205.07069
-
Heavy-Tail Phenomenon in Decentralized SGD 13 May 2022 · 0 repositories · arXiv:2205.06689
-
A Communication-Efficient Distributed Gradient Clipping Algorithm for Training Deep Neural Networks 10 May 2022 · 1 repository · arXiv:2205.05040Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Unsupervised Slot Schema Induction for Task-oriented Dialog 9 May 2022 · 0 repositories · arXiv:2205.04515
-
Federated Random Reshuffling with Compression and Variance Reduction 8 May 2022 · 0 repositories · arXiv:2205.03914
-
Making SGD Parameter-Free 4 May 2022 · 0 repositories · arXiv:2205.02160
-
The Directional Bias Helps Stochastic Gradient Descent to Generalize in Kernel Regression Models 29 Apr 2022 · 0 repositories · arXiv:2205.00061
-
Beyond Lipschitz: Sharp Generalization and Excess Risk Bounds for Full-Batch GD 26 Apr 2022 · 0 repositories · arXiv:2204.12446
-
Federated Stochastic Primal-dual Learning with Differential Privacy 26 Apr 2022 · 0 repositories · arXiv:2204.12284
-
On Acceleration of Gradient-Based Empirical Risk Minimization using Local Polynomial Regression 16 Apr 2022 · 0 repositories · arXiv:2204.07702
-
Dynamic Schema Graph Fusion Network for Multi-Domain Dialogue State Tracking 14 Apr 2022 · 0 repositories · arXiv:2204.06677
-
What You See is What You Get: Principled Deep Learning via Distributional Generalization 7 Apr 2022 · 1 repository · arXiv:2204.03230Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Nonlinear gradient mappings and stochastic optimization: A general framework with applications to heavy-tail noise 6 Apr 2022 · 0 repositories · arXiv:2204.02593
-
Privacy-Preserving Federated Learning via System Immersion and Random Matrix Encryption 5 Apr 2022 · 0 repositories · arXiv:2204.02497
-
Deep learning, stochastic gradient descent and diffusion maps 4 Apr 2022 · 0 repositories · arXiv:2204.01365
-
Data Sampling Affects the Complexity of Online SGD over Dependent Data 31 Mar 2022 · 0 repositories · arXiv:2204.00006
-
Exploiting Explainable Metrics for Augmented SGD 31 Mar 2022 · 2 repositories · arXiv:2203.16723
-
Conjugate Gradient Method for Generative Adversarial Networks 28 Mar 2022 · 1 repository · arXiv:2203.14495
-
A Robust Optimization Method for Label Noisy Datasets Based on Adaptive Threshold: Adaptive-k 26 Mar 2022 · 1 repository · arXiv:2203.14165