Methods › General › Stochastic Optimization › SGD › Papers, page 2
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 2 of 21: papers 101 to 200 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Algorithmic Stability of Stochastic Gradient Descent with Momentum under Heavy-Tailed Noise 2 Feb 2025 · 0 repositories · arXiv:2502.00885
-
Understanding Why Adam Outperforms SGD: Gradient Heterogeneity in Transformers 31 Jan 2025 · 1 repository · arXiv:2502.00213
-
A Unified Analysis of Stochastic Gradient Descent with Arbitrary Data Permutations and Beyond 27 Jan 2025 · 0 repositories · arXiv:2501.16117
-
Ringmaster ASGD: The First Asynchronous SGD with Optimal Time Complexity 27 Jan 2025 · 0 repositories · arXiv:2501.16168
-
Mathematical analysis of the gradients in deep learning 26 Jan 2025 · 0 repositories · arXiv:2501.15646
-
Scalable Decentralized Learning with Teleportation 25 Jan 2025 · 0 repositories · arXiv:2501.15259
-
A Near-optimal Algorithm for Learning Margin Halfspaces with Massart Noise 16 Jan 2025 · 0 repositories · arXiv:2501.09691
-
Identification of Traditional Medicinal Plant Leaves Using an effective Deep Learning model and Self-Curated Dataset 16 Jan 2025 · 0 repositories · arXiv:2501.09363
-
Increasing Batch Size Improves Convergence of Stochastic Gradient Descent with Momentum 15 Jan 2025 · 0 repositories · arXiv:2501.08883
-
Is Stochastic Gradient Descent Effective? A PDE Perspective on Machine Learning processes 14 Jan 2025 · 0 repositories · arXiv:2501.08425
-
Communication-Efficient, 2D Parallel Stochastic Gradient Descent for Distributed-Memory Optimization 13 Jan 2025 · 0 repositories · arXiv:2501.07526
-
Averaged Adam accelerates stochastic optimization in the training of deep neural network approximations for partial differential equation and optimal control problems 10 Jan 2025 · 1 repository · arXiv:2501.06081
-
Revisiting LocalSGD and SCAFFOLD: Improved Rates and Missing Analysis 8 Jan 2025 · 0 repositories · arXiv:2501.04443
-
Mixing Times and Privacy Analysis for the Projected Langevin Algorithm under a Modulus of Continuity 7 Jan 2025 · 0 repositories · arXiv:2501.04134
-
ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think 2 Jan 2025 · 0 repositories · arXiv:2501.01045
-
Investigating the Role of Weight Decay in Enhancing Nonconvex SGD 1 Jan 2025 · 0 repositories
-
Accelerating Energy-Efficient Federated Learning in Cell-Free Networks with Adaptive Quantization 30 Dec 2024 · 0 repositories · arXiv:2412.20785
-
ProtScan: Modeling and Prediction of RNA-Protein Interactions 30 Dec 2024 · 0 repositories · arXiv:2412.20933
-
Edge of Stochastic Stability: Revisiting the Edge of Stability for SGD 29 Dec 2024 · 0 repositories · arXiv:2412.20553
-
Self-Assembly of a Biologically Plausible Learning Circuit 28 Dec 2024 · 0 repositories · arXiv:2412.20018
-
On the Convergence of DP-SGD with Adaptive Clipping 27 Dec 2024 · 0 repositories · arXiv:2412.19916
-
Torque-Aware Momentum 25 Dec 2024 · 0 repositories · arXiv:2412.18790
-
Learning to Generate Gradients for Test-Time Adaptation via Test-Time Training Layers 22 Dec 2024 · 1 repository · arXiv:2412.16901
-
Gradient-Based Non-Linear Inverse Learning 21 Dec 2024 · 0 repositories · arXiv:2412.16794
-
Foxtsage vs. Adam: Revolution or Evolution in Optimization? 20 Dec 2024 · 0 repositories · arXiv:2412.17855
-
SWAN: SGD with Normalization and Whitening Enables Stateless LLM Training 17 Dec 2024 · 0 repositories · arXiv:2412.13148
-
Explicit and Implicit Graduated Optimization in Deep Neural Networks 16 Dec 2024 · 1 repository · arXiv:2412.11501
-
No More Adam: Learning Rate Scaling at Initialization is All You Need 16 Dec 2024 · 1 repository · arXiv:2412.11768
-
Coupling-based Convergence Diagnostic and Stepsize Scheme for Stochastic Gradient Descent 15 Dec 2024 · 1 repository · arXiv:2412.11341
-
Stochastic Gradient Descent in the Optimal Control of Execution Costs 14 Dec 2024 · 0 repositories · arXiv:2412.12199
-
Understand the Effectiveness of Shortcuts through the Lens of DCA 13 Dec 2024 · 0 repositories · arXiv:2412.09853
-
EDiT: A Local-SGD-Based Efficient Distributed Training Method for Large Language Models 10 Dec 2024 · 2 repositories · arXiv:2412.07210
-
Stochastic Gradient Descent Revisited 8 Dec 2024 · 0 repositories · arXiv:2412.06070
-
Adaptive Optimization for Enhanced Efficiency in Large-Scale Language Model Training 6 Dec 2024 · 0 repositories · arXiv:2412.04718
-
A Granger-Causal Perspective on Gradient Descent with Application to Pruning 4 Dec 2024 · 0 repositories · arXiv:2412.03035
-
Machine Learning Methods for Automated Interstellar Object Classification with LSST 3 Dec 2024 · 0 repositories · arXiv:2412.02112
-
Memory-Efficient Training for Deep Speaker Embedding Learning in Speaker Verification 2 Dec 2024 · 0 repositories · arXiv:2412.01195
-
PROFIT: A Specialized Optimizer for Deep Fine Tuning 2 Dec 2024 · 0 repositories · arXiv:2412.01930
-
Training Multi-Layer Binary Neural Networks With Local Binary Error Signals 28 Nov 2024 · 0 repositories · arXiv:2412.00119
-
Exponential Moving Average of Weights in Deep Learning: Dynamics and Benefits 27 Nov 2024 · 0 repositories · arXiv:2411.18704
-
Lion Cub: Minimizing Communication Overhead in Distributed Lion 25 Nov 2024 · 0 repositories · arXiv:2411.16462
-
Adaptive Methods through the Lens of SDEs: Theoretical Insights on the Role of Noise 24 Nov 2024 · 0 repositories · arXiv:2411.15958
-
A Unified Analysis for Finite Weight Averaging 20 Nov 2024 · 0 repositories · arXiv:2411.13169
-
Classification of Geographical Land Structure Using Convolution Neural Network and Transfer Learning 19 Nov 2024 · 0 repositories · arXiv:2411.12415
-
Convergence Rate Analysis of LION 12 Nov 2024 · 0 repositories · arXiv:2411.07724
-
FRUGAL: Memory-Efficient Optimization by Reducing State Overhead for Scalable Training 12 Nov 2024 · 1 repository · arXiv:2411.07837Syntology official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 4 where Syntology's instrument failed) · 0 unverified (of 8 harvested samples) · 2 pointer-only (licence)
-
Lean and Mean Adaptive Optimization via Subset-Norm and Subspace-Momentum with Convergence Guarantees 11 Nov 2024 · 0 repositories · arXiv:2411.07120
-
General framework for online-to-nonconvex conversion: Schedule-free SGD is also effective for nonconvex optimization 11 Nov 2024 · 0 repositories · arXiv:2411.07061
-
Impact of Label Noise on Learning Complex Features 7 Nov 2024 · 0 repositories · arXiv:2411.04569
-
A Bayesian Approach to Data Point Selection 6 Nov 2024 · 0 repositories · arXiv:2411.03768
-
Point processes with event time uncertainty 5 Nov 2024 · 0 repositories · arXiv:2411.02694
-
Learning and Transferring Sparse Contextual Bigrams with Linear Transformers 30 Oct 2024 · 0 repositories · arXiv:2410.23438
-
Developing Convolutional Neural Networks using a Novel Lamarckian Co-Evolutionary Algorithm 29 Oct 2024 · 0 repositories · arXiv:2410.22487
-
Shuffling Gradient-Based Methods for Nonconvex-Concave Minimax Optimization 29 Oct 2024 · 0 repositories · arXiv:2410.22297
-
Trustworthiness of Stochastic Gradient Descent in Distributed Learning 28 Oct 2024 · 0 repositories · arXiv:2410.21491
-
Differentially Private Learning Needs Better Model Initialization and Self-Distillation 23 Oct 2024 · 1 repository · arXiv:2410.17566
-
Gradient Normalization Provably Benefits Nonconvex SGD under Heavy-Tailed Noise 21 Oct 2024 · 0 repositories · arXiv:2410.16561
-
Simplicity Bias via Global Convergence of Sharpness Minimization 21 Oct 2024 · 0 repositories · arXiv:2410.16401
-
Unleashing the Potential of Multi-Channel Fusion in Retrieval for Personalized Recommendations 21 Oct 2024 · 0 repositories · arXiv:2410.16080
-
Beyond Discretization: Learning the Optimal Solution Path 18 Oct 2024 · 1 repository · arXiv:2410.14885
-
Implicit Regularization of Sharpness-Aware Minimization for Scale-Invariant Problems 18 Oct 2024 · 0 repositories · arXiv:2410.14802
-
SGD Jittering: A Training Strategy for Robust and Accurate Model-Based Architectures 18 Oct 2024 · 0 repositories · arXiv:2410.14667
-
Active-Dormant Attention Heads: Mechanistically Demystifying Extreme-Token Phenomena in LLMs 17 Oct 2024 · 1 repository · arXiv:2410.13835Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Evaluating Self-Generated Documents for Enhancing Retrieval-Augmented Generation with Large Language Models 17 Oct 2024 · 0 repositories · arXiv:2410.13192
-
From Gradient Clipping to Normalization for Heavy Tailed SGD 17 Oct 2024 · 0 repositories · arXiv:2410.13849
-
Mitigating Hallucinations in Large Vision-Language Models via Summary-Guided Decoding 17 Oct 2024 · 1 repository · arXiv:2410.13321
-
Nonlinear Stochastic Gradient Descent and Heavy-tailed Noise: A Unified Framework and High-probability Guarantees 17 Oct 2024 · 0 repositories · arXiv:2410.13954
-
Advanced Persistent Threats (APT) Attribution Using Deep Reinforcement Learning 15 Oct 2024 · 0 repositories · arXiv:2410.11463
-
Age-of-Gradient Updates for Federated Learning over Random Access Channels 15 Oct 2024 · 0 repositories · arXiv:2410.11986
-
Non-convergence to global minimizers in data driven supervised deep learning: Adam and stochastic gradient descent optimization provably fail to converge to global minimizers in the training of deep neural networks with ReLU activation 14 Oct 2024 · 0 repositories · arXiv:2410.10533
-
SAMPa: Sharpness-aware Minimization Parallelized 14 Oct 2024 · 1 repository · arXiv:2410.10683Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Sharpness-Aware Minimization Efficiently Selects Flatter Minima Late in Training 14 Oct 2024 · 0 repositories · arXiv:2410.10373Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Learning Orthogonal Multi-Index Models: A Fine-Grained Information Exponent Analysis 13 Oct 2024 · 0 repositories · arXiv:2410.09678
-
Sharper Guarantees for Learning Neural Network Classifiers with Gradient Methods 13 Oct 2024 · 0 repositories · arXiv:2410.10024
-
Data Deletion for Linear Regression with Noisy SGD 12 Oct 2024 · 0 repositories · arXiv:2410.09311
-
Zeroth-Order Fine-Tuning of LLMs in Random Subspaces 11 Oct 2024 · 1 repository · arXiv:2410.08989Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 5 where Syntology's instrument failed) · 4 unverified (of 15 harvested samples) · 15 pointer-only (licence)
-
Adam Exploits ℓ_∞-geometry of Loss Landscape via Coordinate-wise Adaptivity 10 Oct 2024 · 1 repository · arXiv:2410.08198Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 10 harvested samples)
-
On the Convergence of (Stochastic) Gradient Descent for Kolmogorov--Arnold Networks 10 Oct 2024 · 0 repositories · arXiv:2410.08041
-
A second-order-like optimizer with adaptive gradient scaling for deep learning 8 Oct 2024 · 1 repository · arXiv:2410.05871
-
Asynchronous Stochastic Gradient Descent with Decoupled Backpropagation and Layer-Wise Updates 8 Oct 2024 · 0 repositories · arXiv:2410.05985
-
Unveiling the Backbone-Optimizer Coupling Bias in Visual Representation Learning 8 Oct 2024 · 0 repositories · arXiv:2410.06373
-
Nonasymptotic Analysis of Stochastic Gradient Descent with the Richardson-Romberg Extrapolation 7 Oct 2024 · 0 repositories · arXiv:2410.05106
-
A Comprehensive Framework for Analyzing the Convergence of Adam: Bridging the Gap with SGD 6 Oct 2024 · 0 repositories · arXiv:2410.04458
-
MindFlayer SGD: Efficient Parallel SGD in the Presence of Heterogeneous and Random Worker Compute Times 5 Oct 2024 · 0 repositories · arXiv:2410.04285
-
SGD with memory: fundamental properties and stochastic acceleration 5 Oct 2024 · 0 repositories · arXiv:2410.04228
-
Collaborative and Efficient Personalization with Mixtures of Adaptors 4 Oct 2024 · 0 repositories · arXiv:2410.03497
-
Estimating Generalization Performance Along the Trajectory of Proximal SGD in Robust Regression 3 Oct 2024 · 1 repository · arXiv:2410.02629
-
Towards Better Generalization: Weight Decay Induces Low-rank Bias for Neural Networks 3 Oct 2024 · 0 repositories · arXiv:2410.02176
-
Stochastic Gradient Descent with Adaptive Data 2 Oct 2024 · 0 repositories · arXiv:2410.01195
-
Truncated Kernel Stochastic Gradient Descent on Spheres 2 Oct 2024 · 0 repositories · arXiv:2410.01570
-
Comparing Unidirectional, Bidirectional, and Word2vec Models for Discovering Vulnerabilities in Compiled Lifted Code 26 Sep 2024 · 0 repositories · arXiv:2409.17513
-
Does Worst-Performing Agent Lead the Pack? Analyzing Agent Dynamics in Unified Distributed SGD 26 Sep 2024 · 0 repositories · arXiv:2409.17499
-
Differential Privacy Regularization: Protecting Training Data Through Loss Function Regularization 25 Sep 2024 · 0 repositories · arXiv:2409.17144
-
Convergence of Distributed Adaptive Optimization with Local Updates 20 Sep 2024 · 0 repositories · arXiv:2409.13155
-
Exploring Scaling Laws for Local SGD in Large Language Model Training 20 Sep 2024 · 0 repositories · arXiv:2409.13198
-
The Optimality of (Accelerated) SGD for High-Dimensional Quadratic Optimization 15 Sep 2024 · 0 repositories · arXiv:2409.09745
-
A Dynamic Weighting Strategy to Mitigate Worker Node Failure in Distributed Deep Learning 14 Sep 2024 · 0 repositories · arXiv:2409.09242
-
Increasing Both Batch Size and Learning Rate Accelerates Stochastic Gradient Descent 13 Sep 2024 · 1 repository · arXiv:2409.08770
-
Asymptotics of Stochastic Gradient Descent with Dropout Regularization in Linear Models 11 Sep 2024 · 1 repository · arXiv:2409.07434
-
Enhancing Deep Learning with Optimized Gradient Descent: Bridging Numerical Methods and Neural Network Training 7 Sep 2024 · 0 repositories · arXiv:2409.04707