Methods › General › Stochastic Optimization › SGD › Papers, page 6
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 6 of 21: papers 501 to 600 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning 7 Jun 2023 · 1 repository · arXiv:2306.04815
-
Prompter: Zero-shot Adaptive Prefixes for Dialogue State Tracking Domain Adaptation 7 Jun 2023 · 1 repository · arXiv:2306.04724
-
Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler Subnetworks 7 Jun 2023 · 1 repository · arXiv:2306.04251Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Machine learning in and out of equilibrium 6 Jun 2023 · 1 repository · arXiv:2306.03521
-
Decentralized SGD and Average-direction SAM are Asymptotically Equivalent 5 Jun 2023 · 1 repository · arXiv:2306.02913Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples) · 2 pointer-only (licence)
-
Improved Stability and Generalization Guarantees of the Decentralized SGD Algorithm 5 Jun 2023 · 0 repositories · arXiv:2306.02939
-
Online Bootstrap Inference with Nonconvex Stochastic Gradient Descent Estimator 3 Jun 2023 · 0 repositories · arXiv:2306.02205
-
Towards Sustainable Learning: Coresets for Data-efficient Deep Learning 2 Jun 2023 · 1 repository · arXiv:2306.01244Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Improving Energy Conserving Descent for Machine Learning: Theory and Practice 1 Jun 2023 · 1 repository · arXiv:2306.00352
-
A Bayesian Approach To Analysing Training Data Attribution In Deep Learning 31 May 2023 · 1 repository · arXiv:2305.19765Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 5 harvested samples)
-
The Impact of Positional Encoding on Length Generalization in Transformers 31 May 2023 · 2 repositories · arXiv:2305.19466Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Toward Understanding Why Adam Converges Faster Than SGD for Transformers 31 May 2023 · 0 repositories · arXiv:2306.00204
-
BiSLS/SPS: Auto-tune Step Sizes for Stable Bi-level Optimization 30 May 2023 · 0 repositories · arXiv:2305.18666
-
On Convergence of Incremental Gradient for Non-Convex Smooth Functions 30 May 2023 · 0 repositories · arXiv:2305.19259
-
A Rainbow in Deep Network Black Boxes 29 May 2023 · 1 repository · arXiv:2305.18512Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 14 harvested samples) · 1 pointer-only (licence)
-
Convergence of AdaGrad for Non-convex Objectives: Simple Proofs and Relaxed Assumptions 29 May 2023 · 0 repositories · arXiv:2305.18471
-
Escaping mediocrity: how two-layer networks learn hard generalized linear models with SGD 29 May 2023 · 2 repositories · arXiv:2305.18502
-
Acceleration of stochastic gradient descent with momentum by averaging: finite-sample rates and asymptotic normality 28 May 2023 · 0 repositories · arXiv:2305.17665
-
Direction-oriented Multi-objective Learning: Simple and Provable Stochastic Algorithms 28 May 2023 · 1 repository · arXiv:2305.18409Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 1 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 10 harvested samples)
-
The Implicit Regularization of Dynamical Stability in Stochastic Gradient Descent 27 May 2023 · 0 repositories · arXiv:2305.17490
-
Rotational Equilibrium: How Weight Decay Balances Learning Across Neural Networks 26 May 2023 · 2 repositories · arXiv:2305.17212Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Schema-Guided User Satisfaction Modeling for Task-Oriented Dialogues 26 May 2023 · 1 repository · arXiv:2305.16798
-
XGrad: Boosting Gradient-Based Optimizers With Weight Prediction 26 May 2023 · 1 repository · arXiv:2305.18240
-
ADLER -- An efficient Hessian-based strategy for adaptive learning rate 25 May 2023 · 0 repositories · arXiv:2305.16396
-
Exploiting Noise as a Resource for Computation and Learning in Spiking Neural Networks 25 May 2023 · 1 repository · arXiv:2305.16044Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Implicit bias of SGD in L₂-regularized linear DNNs: One-way jumps from high to low rank 25 May 2023 · 0 repositories · arXiv:2305.16038
-
Incentivizing Honesty among Competitors in Collaborative Learning and Optimization 25 May 2023 · 0 repositories · arXiv:2305.16272
-
Neural Characteristic Activation Analysis and Geometric Parameterization for ReLU Networks 25 May 2023 · 1 repository · arXiv:2305.15912Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 5 harvested samples)
-
Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer 25 May 2023 · 0 repositories · arXiv:2305.16380
-
Utility-Probability Duality of Neural Networks 24 May 2023 · 0 repositories · arXiv:2305.14859
-
Local SGD Accelerates Convergence by Exploiting Second Order Information of the Loss Function 24 May 2023 · 0 repositories · arXiv:2305.15013
-
On the Convergence of Black-Box Variational Inference 24 May 2023 · 0 repositories · arXiv:2305.15349
-
Layer-wise Adaptive Step-Sizes for Stochastic First-Order Methods for Deep Learning 23 May 2023 · 0 repositories · arXiv:2305.13664
-
Two Sides of One Coin: the Limits of Untuned SGD and the Power of Adaptive Methods 21 May 2023 · 0 repositories · arXiv:2305.12475
-
Evolutionary Algorithms in the Light of SGD: Limit Equivalence, Minima Flatness, and Transfer Learning 20 May 2023 · 0 repositories · arXiv:2306.09991
-
GraVAC: Adaptive Compression for Communication-Efficient Distributed DL Training 20 May 2023 · 1 repository · arXiv:2305.12201
-
Stability and Generalization of lp-Regularized Stochastic Learning for GCN 20 May 2023 · 0 repositories · arXiv:2305.12085
-
Uniform-in-Time Wasserstein Stability Bounds for (Noisy) Stochastic Gradient Descent 20 May 2023 · 0 repositories · arXiv:2305.12056
-
Conditional Online Learning for Keyword Spotting 19 May 2023 · 0 repositories · arXiv:2305.13332
-
Smoothing the Landscape Boosts the Signal for SGD: Optimal Sample Complexity for Learning Single Index Models 18 May 2023 · 0 repositories · arXiv:2305.10633
-
Stochastic Ratios Tracking Algorithm for Large Scale Machine Learning Problems 17 May 2023 · 0 repositories · arXiv:2305.09978
-
Faster Federated Learning with Decaying Number of Local SGD Steps 16 May 2023 · 0 repositories · arXiv:2305.09628
-
MoMo: Momentum Models for Adaptive Learning Rates 12 May 2023 · 1 repository · arXiv:2305.07583Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Online Learning Under A Separable Stochastic Approximation Framework 12 May 2023 · 0 repositories · arXiv:2305.07484
-
Securing Distributed SGD against Gradient Leakage Threats 10 May 2023 · 1 repository · arXiv:2305.06473
-
FedHB: Hierarchical Bayesian Federated Learning 8 May 2023 · 0 repositories · arXiv:2305.04979
-
Semi-Asynchronous Federated Edge Learning Mechanism via Over-the-air Computation 6 May 2023 · 0 repositories · arXiv:2305.04066
-
A Momentum-Incorporated Non-Negative Latent Factorization of Tensors Model for Dynamic Network Representation 4 May 2023 · 0 repositories · arXiv:2305.02782
-
Revisiting Gradient Clipping: Stochastic bias and tight convergence guarantees 2 May 2023 · 0 repositories · arXiv:2305.01588
-
Fairness Uncertainty Quantification: How certain are you that the model is fair? 27 Apr 2023 · 0 repositories · arXiv:2304.13950
-
Noise Is Not the Main Factor Behind the Gap Between SGD and Adam on Transformers, but Sign Descent Might Be 27 Apr 2023 · 1 repository · arXiv:2304.13960
-
Hierarchical Weight Averaging for Deep Neural Networks 23 Apr 2023 · 0 repositories · arXiv:2304.11519
-
An Incomplete Tensor Tucker decomposition based Traffic Speed Prediction Method 21 Apr 2023 · 0 repositories · arXiv:2304.10961
-
Loss Minimization Yields Multicalibration for Large Neural Networks 19 Apr 2023 · 0 repositories · arXiv:2304.09424
-
Convergence of stochastic gradient descent under a local Lojasiewicz condition for deep neural networks 18 Apr 2023 · 0 repositories · arXiv:2304.09221
-
Fast and Straggler-Tolerant Distributed SGD with Reduced Computation Load 17 Apr 2023 · 0 repositories · arXiv:2304.08589
-
Landslide Susceptibility Prediction Modeling Based on Self-Screening Deep Learning Model 12 Apr 2023 · 0 repositories · arXiv:2304.06054
-
High-dimensional scaling limits and fluctuations of online least-squares SGD with smooth covariance 3 Apr 2023 · 0 repositories · arXiv:2304.00707
-
Fast Convergence of Random Reshuffling under Over-Parameterization and the Polyak-Łojasiewicz Condition 2 Apr 2023 · 0 repositories · arXiv:2304.00459
-
Doubly Stochastic Models: Learning with Unbiased Label Noises and Inference Stability 1 Apr 2023 · 0 repositories · arXiv:2304.00320
-
On the Query Complexity of Training Data Reconstruction in Private Learning 29 Mar 2023 · 0 repositories · arXiv:2303.16372
-
Zero-Shot Generalizable End-to-End Task-Oriented Dialog System using Context Summarization and Domain Schema 28 Mar 2023 · 1 repository · arXiv:2303.16252
-
Toward Open-domain Slot Filling via Self-supervised Co-training 24 Mar 2023 · 0 repositories · arXiv:2303.13801
-
Type-II Saddles and Probabilistic Stability of Stochastic Gradient Descent 23 Mar 2023 · 0 repositories · arXiv:2303.13093
-
𝒞ᵏ-continuous Spline Approximation with TensorFlow Gradient Descent Optimizers 22 Mar 2023 · 0 repositories · arXiv:2303.12454
-
Lower Generalization Bounds for GD and SGD in Smooth Stochastic Convex Optimization 19 Mar 2023 · 0 repositories · arXiv:2303.10758
-
Convergence Analysis of Stochastic Gradient Descent with MCMC Estimators 19 Mar 2023 · 0 repositories · arXiv:2303.10599
-
Practical and Matching Gradient Variance Bounds for Black-Box Variational Bayesian Inference 18 Mar 2023 · 0 repositories · arXiv:2303.10472
-
Tighter Lower Bounds for Shuffling SGD: Random Permutations and Beyond 13 Mar 2023 · 1 repository · arXiv:2303.07160Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Scavenger: A Cloud Service for Optimizing Cost and Performance of ML Training 12 Mar 2023 · 0 repositories · arXiv:2303.06659
-
Fast Latent Factor Analysis via a Fuzzy PID-Incorporated Stochastic Gradient Descent Algorithm 7 Mar 2023 · 0 repositories · arXiv:2303.03941
-
Revisiting the Noise Model of Stochastic Gradient Descent 5 Mar 2023 · 0 repositories · arXiv:2303.02749
-
Gradient Norm Aware Minimization Seeks First-Order Flatness and Improves Generalization 3 Mar 2023 · 1 repository · arXiv:2303.03108
-
Implicit Stochastic Gradient Descent for Training Physics-informed Neural Networks 3 Mar 2023 · 0 repositories · arXiv:2303.01767
-
Finite-Sample Analysis of Learning High-Dimensional Single ReLU Neuron 3 Mar 2023 · 0 repositories · arXiv:2303.02255
-
Dropout Reduces Underfitting 2 Mar 2023 · 1 repository · arXiv:2303.01500Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Over-training with Mixup May Hurt Generalization 2 Mar 2023 · 0 repositories · arXiv:2303.01475
-
Variance-reduced Clipping for Non-convex Optimization 2 Mar 2023 · 1 repository · arXiv:2303.00883Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
Why (and When) does Local SGD Generalize Better than SGD? 2 Mar 2023 · 1 repository · arXiv:2303.01215Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples)
-
AdaSAM: Boosting Sharpness-Aware Minimization with Adaptive Learning Rate and Momentum for Training Deep Neural Networks 1 Mar 2023 · 0 repositories · arXiv:2303.00565
-
An Efficient Tester-Learner for Halfspaces 28 Feb 2023 · 0 repositories · arXiv:2302.14853
-
Fast as CHITA: Neural Network Pruning with Combinatorial Optimization 28 Feb 2023 · 0 repositories · arXiv:2302.14623
-
High Probability Convergence of Stochastic Gradient Methods 28 Feb 2023 · 0 repositories · arXiv:2302.14843
-
Stochastic Gradient Descent under Markovian Sampling Schemes 28 Feb 2023 · 0 repositories · arXiv:2302.14428
-
A Unified Framework for Soft Threshold Pruning 25 Feb 2023 · 1 repository · arXiv:2302.13019Syntology official (archive's flag): 2 ran · 2 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 1 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
On the Training Instability of Shuffling SGD with Batch Normalization 24 Feb 2023 · 0 repositories · arXiv:2302.12444
-
Statistical Inference with Stochastic Gradient Methods under φ-mixing Data 24 Feb 2023 · 0 repositories · arXiv:2302.12717
-
Random Teachers are Good Teachers 23 Feb 2023 · 1 repository · arXiv:2302.12091
-
Multi-Message Shuffled Privacy in Federated Learning 22 Feb 2023 · 0 repositories · arXiv:2302.11152
-
SGD learning on neural networks: leap complexity and saddle-to-saddle dynamics 21 Feb 2023 · 0 repositories · arXiv:2302.11055
-
Statistical Inference for Linear Functionals of Online SGD in High-dimensional Linear Regression 20 Feb 2023 · 0 repositories · arXiv:2302.09727
-
mSAM: Micro-Batch-Averaged Sharpness-Aware Minimization 19 Feb 2023 · 0 repositories · arXiv:2302.09693
-
The Generalization Error of Stochastic Mirror Descent on Over-Parametrized Linear Models 18 Feb 2023 · 0 repositories · arXiv:2302.09433
-
(S)GD over Diagonal Linear Networks: Implicit Regularisation, Large Stepsizes and Edge of Stability 17 Feb 2023 · 0 repositories · arXiv:2302.08982
-
FOSI: Hybrid First and Second Order Optimization 16 Feb 2023 · 1 repository · arXiv:2302.08484Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Almost Sure Saddle Avoidance of Stochastic Gradient Methods without the Bounded Gradient Assumption 15 Feb 2023 · 0 repositories · arXiv:2302.07862
-
From high-dimensional & mean-field dynamics to dimensionless ODEs: A unifying approach to SGD in two-layers networks 12 Feb 2023 · 1 repository · arXiv:2302.05882
-
Cyclic and Randomized Stepsizes Invoke Heavier Tails in SGD than Constant Stepsize 10 Feb 2023 · 0 repositories · arXiv:2302.05516
-
On the Convergence of Stochastic Gradient Descent for Linear Inverse Problems in Banach Spaces 10 Feb 2023 · 0 repositories · arXiv:2302.05197
-
Information Theoretic Lower Bounds for Information Theoretic Upper Bounds 9 Feb 2023 · 0 repositories · arXiv:2302.04925