Methods › General › Stochastic Optimization › SGD › Papers, page 3
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 3 of 21: papers 201 to 300 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Fast Forwarding Low-Rank Training 6 Sep 2024 · 0 repositories · arXiv:2409.04206
-
Bootstrap SGD: Algorithmic Stability and Robustness 2 Sep 2024 · 0 repositories · arXiv:2409.01074
-
1-Bit FQT: Pushing the Limit of Fully Quantized Training to 1-bit 26 Aug 2024 · 1 repository · arXiv:2408.14267
-
PFDiff: Training-free Acceleration of Diffusion Models through the Gradient Guidance of Past and Future 16 Aug 2024 · 1 repository · arXiv:2408.08822Syntology official (archive's flag): 7 ran · 8 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 6 where Syntology's instrument failed) · 1 unverified (of 9 harvested samples) · 4 pointer-only (licence)
-
Langevin dynamics for high-dimensional optimization: the case of multi-spiked tensor PCA 12 Aug 2024 · 0 repositories · arXiv:2408.06401
-
Incremental Gauss-Newton Descent for Machine Learning 10 Aug 2024 · 1 repository · arXiv:2408.05560
-
Optimization-Driven Adaptive Experimentation 8 Aug 2024 · 0 repositories · arXiv:2408.04570
-
On Using Quasirandom Sequences in Machine Learning for Model Weight Initialization 5 Aug 2024 · 1 repository · arXiv:2408.02654
-
Optimizing Cox Models with Stochastic Gradient Descent: Theoretical Foundations and Practical Guidances 5 Aug 2024 · 0 repositories · arXiv:2408.02839
-
Transforming Slot Schema Induction with Generative Dialogue State Inference 3 Aug 2024 · 1 repository · arXiv:2408.01638
-
Can LLMs predict the convergence of Stochastic Gradient Descent? 3 Aug 2024 · 0 repositories · arXiv:2408.01736
-
Data Debugging is NP-hard for Classifiers Trained with SGD 2 Aug 2024 · 0 repositories · arXiv:2408.01365
-
Zero-Shot Cross-Domain Dialogue State Tracking via Dual Low-Rank Adaptation 31 Jul 2024 · 1 repository · arXiv:2407.21633Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 1 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
No learning rates needed: Introducing SALSA -- Stable Armijo Line Search Adaptation 30 Jul 2024 · 1 repository · arXiv:2407.20650
-
The Entrapment Problem in Random Walk Decentralized Learning 30 Jul 2024 · 1 repository · arXiv:2407.20611
-
Convergence rates for the Adam optimizer 29 Jul 2024 · 0 repositories · arXiv:2407.21078
-
Characterizing Dynamical Stability of Stochastic Gradient Descent in Overparameterized Learning 29 Jul 2024 · 0 repositories · arXiv:2407.20209
-
Ordered Momentum for Asynchronous SGD 27 Jul 2024 · 0 repositories · arXiv:2407.19234
-
A New Theoretical Perspective on Data Heterogeneity in Federated Optimization 22 Jul 2024 · 0 repositories · arXiv:2407.15567
-
Training Zero-Shot Generalizable End-to-End Task-Oriented Dialog System Without Turn-level Dialog Annotations 21 Jul 2024 · 0 repositories · arXiv:2407.15055
-
Non-convergence of Adam and other adaptive stochastic gradient descent optimization methods for non-vanishing learning rates 11 Jul 2024 · 0 repositories · arXiv:2407.08100
-
Predicting Heart Failure with Attention Learning Techniques Utilizing Cardiovascular Data 11 Jul 2024 · 0 repositories · arXiv:2407.08289
-
Deconstructing What Makes a Good Optimizer for Language Models 10 Jul 2024 · 0 repositories · arXiv:2407.07972
-
Stochastic Gradient Descent for Two-layer Neural Networks 10 Jul 2024 · 0 repositories · arXiv:2407.07670
-
LoCo: Low-Bit Communication Adaptor for Large-scale Model Training 5 Jul 2024 · 1 repository · arXiv:2407.04480
-
Bias of Stochastic Gradient Descent or the Architecture: Disentangling the Effects of Overparameterization of Neural Networks 4 Jul 2024 · 0 repositories · arXiv:2407.03848
-
Gradient descent with generalized Newton's method 3 Jul 2024 · 1 repository · arXiv:2407.02772Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Universal Length Generalization with Turing Programs 3 Jul 2024 · 0 repositories · arXiv:2407.03310
-
Stochastic Differential Equations models for Least-Squares Stochastic Gradient Descent 2 Jul 2024 · 0 repositories · arXiv:2407.02322
-
Infinite Width Models That Work: Why Feature Learning Doesn't Matter as Much as You Think 27 Jun 2024 · 0 repositories · arXiv:2406.18800
-
Towards Efficient and Scalable Training of Differentially Private Deep Learning 25 Jun 2024 · 1 repository · arXiv:2406.17298
-
Effect of Random Learning Rate: Theoretical Analysis of SGD Dynamics in Non-Convex Optimization via Stationary Distribution 23 Jun 2024 · 0 repositories · arXiv:2406.16032
-
AdaGrad under Anisotropic Smoothness 21 Jun 2024 · 0 repositories · arXiv:2406.15244
-
Communication-Efficient Adaptive Batch Size Strategies for Distributed Local Gradient Methods 20 Jun 2024 · 0 repositories · arXiv:2406.13936
-
Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses 20 Jun 2024 · 1 repository · arXiv:2406.14340
-
Towards Exact Gradient-based Training on Analog In-memory Computing 18 Jun 2024 · 0 repositories · arXiv:2406.12774
-
To Clip or not to Clip: the Dynamics of SGD with Gradient Clipping in High-Dimensions 17 Jun 2024 · 1 repository · arXiv:2406.11733
-
Distributed Stochastic Gradient Descent with Staleness: A Stochastic Delay Differential Equation Based Framework 17 Jun 2024 · 0 repositories · arXiv:2406.11159
-
How Neural Networks Learn the Support is an Implicit Regularization Effect of SGD 17 Jun 2024 · 0 repositories · arXiv:2406.11110
-
Just How Flexible are Neural Networks in Practice? 17 Jun 2024 · 0 repositories · arXiv:2406.11463
-
What is the long-run distribution of stochastic gradient descent? A large deviations analysis 13 Jun 2024 · 0 repositories · arXiv:2406.09241
-
Why Warmup the Learning Rate? Underlying Mechanisms and Improvements 13 Jun 2024 · 0 repositories · arXiv:2406.09405Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
Scaling Laws in Linear Regression: Compute, Parameters, and Data 12 Jun 2024 · 0 repositories · arXiv:2406.08466
-
Provable Complexity Improvement of AdaGrad over SGD: Upper and Lower Bounds in Stochastic Non-Convex Optimization 7 Jun 2024 · 0 repositories · arXiv:2406.04592
-
The Expanding Scope of the Stability Gap: Unveiling its Presence in Joint Incremental Learning of Homogeneous Tasks 7 Jun 2024 · 0 repositories · arXiv:2406.05114
-
Concurrent Training and Layer Pruning of Deep Neural Networks 6 Jun 2024 · 0 repositories · arXiv:2406.04549
-
Stochastic Polyak Step-sizes and Momentum: Convergence Guarantees and Practical Performance 6 Jun 2024 · 1 repository · arXiv:2406.04142Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Online Learning and Information Exponents: On The Importance of Batch size, and Time/Complexity Tradeoffs 4 Jun 2024 · 1 repository · arXiv:2406.02157
-
Cohort Squeeze: Beyond a Single Communication Round per Cohort in Cross-Device Federated Learning 3 Jun 2024 · 0 repositories · arXiv:2406.01115
-
Demystifying SGD with Doubly Stochastic Gradients 3 Jun 2024 · 0 repositories · arXiv:2406.00920
-
Local Methods with Adaptivity via Scaling 2 Jun 2024 · 0 repositories · arXiv:2406.00846
-
Machine Learning Methods for Pricing Financial Derivatives 1 Jun 2024 · 0 repositories · arXiv:2406.00459
-
Stochastic Resetting Mitigates Latent Gradient Bias of SGD from Label Noise 1 Jun 2024 · 0 repositories · arXiv:2406.00396
-
PixOOD: Pixel-Level Out-of-Distribution Detection 30 May 2024 · 1 repository · arXiv:2405.19882
-
Sharpness-Aware Minimization Enhances Feature Quality via Balanced Learning 30 May 2024 · 0 repositories · arXiv:2405.20439
-
Symmetries in Overparametrized Neural Networks: A Mean-Field View 30 May 2024 · 0 repositories · arXiv:2405.19995
-
The High Line: Exact Risk and Learning Rate Curves of Stochastic Adaptive Learning Rate Algorithms 30 May 2024 · 1 repository · arXiv:2405.19585
-
SPABA: A Single-Loop and Probabilistic Stochastic Bilevel Algorithm Achieving Optimal Sample Complexity 29 May 2024 · 0 repositories · arXiv:2405.18777
-
A Hessian-Aware Stochastic Differential Equation for Modelling SGD 28 May 2024 · 0 repositories · arXiv:2405.18373
-
A Margin-based Multiclass Generalization Bound via Geometric Complexity 28 May 2024 · 0 repositories · arXiv:2405.18590
-
The Unified Balance Theory of Second-Moment Exponential Scaling Optimizers in Visual Tasks 28 May 2024 · 0 repositories · arXiv:2405.18498
-
Towards Communication-efficient Federated Learning via Sparse and Aligned Adaptive Optimization 28 May 2024 · 0 repositories · arXiv:2405.17932
-
Dual-Delayed Asynchronous SGD for Arbitrarily Heterogeneous Data 27 May 2024 · 0 repositories · arXiv:2405.16966
-
How Do the Architecture and Optimizer Affect Representation Learning? On the Training Dynamics of Representations in Deep Neural Networks 27 May 2024 · 0 repositories · arXiv:2405.17377
-
AdaFisher: Adaptive Second Order Optimization via Fisher Information 26 May 2024 · 1 repository · arXiv:2405.16397Syntology official (archive's flag): 7 ran · 7 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 5 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Does SGD really happen in tiny subspaces? 25 May 2024 · 0 repositories · arXiv:2405.16002Syntology 14 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Derivatives of Stochastic Gradient Descent in parametric optimization 24 May 2024 · 0 repositories · arXiv:2405.15894
-
Freya PAGE: First Optimal Time Complexity for Large-Scale Nonconvex Finite-Sum Optimization with Heterogeneous Asynchronous Computations 24 May 2024 · 0 repositories · arXiv:2405.15545
-
Exact Gauss-Newton Optimization for Training Deep Neural Networks 23 May 2024 · 1 repository · arXiv:2405.14402
-
Semi-Discrete Optimal Transport: Nearly Minimax Estimation With Stochastic Gradient Descent and Adaptive Entropic Regularization 23 May 2024 · 0 repositories · arXiv:2405.14459
-
Surge Phenomenon in Optimal Learning Rate and Batch Size Scaling 23 May 2024 · 0 repositories · arXiv:2405.14578
-
Score-based Generative Models with Adaptive Momentum 22 May 2024 · 0 repositories · arXiv:2405.13726
-
Uncertainty quantification by block bootstrap for differentially private stochastic gradient descent 21 May 2024 · 0 repositories · arXiv:2405.12553
-
The Limits and Potentials of Local SGD for Distributed Heterogeneous Learning with Intermittent Communication 19 May 2024 · 0 repositories · arXiv:2405.11667
-
Audio-Visual Speech Recognition based on Regulated Transformer and Spatio-Temporal Fusion Strategy for Driver Assistive Systems 9 May 2024 · 1 repository
-
Custom Gradient Estimators are Straight-Through Estimators in Disguise 8 May 2024 · 0 repositories · arXiv:2405.05171
-
One nose but two nostrils: Learn to align with sparse connections between two olfactory cortices 6 May 2024 · 0 repositories · arXiv:2405.03602
-
The Privacy Power of Correlated Noise in Decentralized Learning 2 May 2024 · 1 repository · arXiv:2405.01031Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Q-Newton: Hybrid Quantum-Classical Scheduling for Accelerating Neural Network Training with Newton's Gradient Descent 30 Apr 2024 · 1 repository · arXiv:2405.00252
-
Matching the Statistical Query Lower Bound for k-Sparse Parity Problems with Sign Stochastic Gradient Descent 18 Apr 2024 · 0 repositories · arXiv:2404.12376
-
Singular-limit analysis of gradient descent with noise injection 18 Apr 2024 · 0 repositories · arXiv:2404.12293
-
Clipped SGD Algorithms for Performative Prediction: Tight Bounds for Clipping Bias and Remedies 17 Apr 2024 · 0 repositories · arXiv:2404.10995
-
Developing Lagrangian-based Methods for Nonsmooth Nonconvex Optimization 15 Apr 2024 · 0 repositories · arXiv:2404.09438
-
Proof-of-Learning with Incentive Security 13 Apr 2024 · 0 repositories · arXiv:2404.09005
-
MoPE: Mixture of Prefix Experts for Zero-Shot Dialogue State Tracking 12 Apr 2024 · 1 repository · arXiv:2404.08559
-
Sliding down the stairs: how correlated latent variables accelerate learning with neural networks 12 Apr 2024 · 0 repositories · arXiv:2404.08602
-
Variance-reduced Zeroth-Order Methods for Fine-Tuning Language Models 11 Apr 2024 · 0 repositories · arXiv:2404.08080
-
PIM-Opt: Demystifying Distributed Optimization Algorithms on a Real-World Processing-In-Memory System 10 Apr 2024 · 1 repository · arXiv:2404.07164
-
Simultaneous linear connectivity of neural networks modulo permutation 9 Apr 2024 · 0 repositories · arXiv:2404.06498
-
Variational Stochastic Gradient Descent for Deep Neural Networks 9 Apr 2024 · 1 repository · arXiv:2404.06549
-
On the Convergence of Continual Learning with Adaptive Methods 8 Apr 2024 · 0 repositories · arXiv:2404.05555
-
Trustless Audits without Revealing Data or Models 6 Apr 2024 · 0 repositories · arXiv:2404.04500
-
Faster Convergence of Stochastic Accelerated Gradient Descent under Interpolation 3 Apr 2024 · 0 repositories · arXiv:2404.02378
-
Linear Combination of Saved Checkpoints Makes Consistency and Diffusion Models Better 2 Apr 2024 · 1 repository · arXiv:2404.02241
-
Make Continual Learning Stronger via C-Flat 1 Apr 2024 · 1 repository · arXiv:2404.00986Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 12 harvested samples) · 12 pointer-only (licence)
-
New logarithmic step size for stochastic gradient descent 1 Apr 2024 · 0 repositories · arXiv:2404.01257
-
TablePuppet: A Generic Framework for Relational Federated Learning 23 Mar 2024 · 1 repository · arXiv:2403.15839
-
Differentially Private Next-Token Prediction of Large Language Models 22 Mar 2024 · 1 repository · arXiv:2403.15638Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples)
-
Analysing heavy-tail properties of Stochastic Gradient Descent by means of Stochastic Recurrence Equations 20 Mar 2024 · 0 repositories · arXiv:2403.13868
-
Context-aware LLM-based Safe Control Against Latent Risks 18 Mar 2024 · 0 repositories · arXiv:2403.11863