Methods › General › Stochastic Optimization › SGD › Papers, page 5
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 5 of 21: papers 401 to 500 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
A Quadratic Synchronization Rule for Distributed Deep Learning 22 Oct 2023 · 1 repository · arXiv:2310.14423Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Towards Hyperparameter-Agnostic DNN Training via Dynamical System Insights 21 Oct 2023 · 0 repositories · arXiv:2310.13901
-
Demystifying the Myths and Legends of Nonconvex Convergence of SGD 19 Oct 2023 · 0 repositories · arXiv:2310.12969
-
LASER: Linear Compression in Wireless Distributed Optimization 19 Oct 2023 · 0 repositories · arXiv:2310.13033
-
Jorge: Approximate Preconditioning for GPU-efficient Second-order Optimization 18 Oct 2023 · 0 repositories · arXiv:2310.12298
-
Learning to Generate Parameters of ConvNets for Unseen Image Data 18 Oct 2023 · 1 repository · arXiv:2310.11862
-
An Automatic Learning Rate Schedule Algorithm for Achieving Faster Convergence and Steeper Descent 17 Oct 2023 · 0 repositories · arXiv:2310.11291
-
Butterfly Effects of SGD Noise: Error Amplification in Behavior Cloning and Autoregression 17 Oct 2023 · 0 repositories · arXiv:2310.11428
-
Resampling Stochastic Gradient Descent Cheaply for Efficient Uncertainty Quantification 17 Oct 2023 · 0 repositories · arXiv:2310.11065
-
Contextual Data Augmentation for Task-Oriented Dialog Systems 16 Oct 2023 · 0 repositories · arXiv:2310.10380
-
Adam-family Methods with Decoupled Weight Decay in Deep Learning 13 Oct 2023 · 0 repositories · arXiv:2310.08858
-
Asynchronous Federated Learning with Incentive Mechanism Based on Contract Theory 10 Oct 2023 · 0 repositories · arXiv:2310.06448
-
Dynamical versus Bayesian Phase Transitions in a Toy Model of Superposition 10 Oct 2023 · 0 repositories · arXiv:2310.06301
-
Augmenting Vision-Based Human Pose Estimation with Rotation Matrix 9 Oct 2023 · 0 repositories · arXiv:2310.06068
-
Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RL 6 Oct 2023 · 2 repositories · arXiv:2310.04411Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 4 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Why Do We Need Weight Decay in Modern Deep Learning? 6 Oct 2023 · 1 repository · arXiv:2310.04415Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Spectral alignment of stochastic gradient descent for high-dimensional classification tasks 4 Oct 2023 · 0 repositories · arXiv:2310.03010
-
Hoeffding's Inequality for Markov Chains under Generalized Concentrability Condition 4 Oct 2023 · 0 repositories · arXiv:2310.02941
-
A simple connection from loss flatness to compressed neural representations 3 Oct 2023 · 0 repositories · arXiv:2310.01770
-
Chunking: Continual Learning is not just about Distribution Shift 3 Oct 2023 · 1 repository · arXiv:2310.02206Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 5 harvested samples) · 5 pointer-only (licence)
-
On the Parallel Complexity of Multilevel Monte Carlo in Stochastic Gradient Descent 3 Oct 2023 · 0 repositories · arXiv:2310.02402
-
Symmetric Single Index Learning 3 Oct 2023 · 0 repositories · arXiv:2310.02117
-
Batch-less stochastic gradient descent for compressive learning of deep regularization for image denoising 2 Oct 2023 · 0 repositories · arXiv:2310.03085
-
Improving Dialogue Management: Quality Datasets vs Models 2 Oct 2023 · 1 repository · arXiv:2310.01339
-
Stability and Generalization for Minibatch SGD and Local SGD 2 Oct 2023 · 0 repositories · arXiv:2310.01139
-
A Theoretical Analysis of Noise Geometry in Stochastic Gradient Descent 1 Oct 2023 · 0 repositories · arXiv:2310.00692
-
On Memorization and Privacy Risks of Sharpness Aware Minimization 30 Sep 2023 · 0 repositories · arXiv:2310.00488Syntology 7 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 1 pointer-only (licence)
-
Robust Stochastic Optimization via Gradient Quantile Clipping 29 Sep 2023 · 0 repositories · arXiv:2309.17316
-
AutoEncoding Tree for City Generation and Applications 27 Sep 2023 · 0 repositories · arXiv:2309.15941
-
Fixing the NTK: From Neural Network Linearizations to Exact Convex Programs 26 Sep 2023 · 0 repositories · arXiv:2309.15096
-
SGD Finds then Tunes Features in Two-Layer Neural Networks with near-Optimal Sample Complexity: A Case Study in the XOR problem 26 Sep 2023 · 0 repositories · arXiv:2309.15111
-
Grounding Description-Driven Dialogue State Trackers with Knowledge-Seeking Turns 23 Sep 2023 · 0 repositories · arXiv:2309.13448
-
A Guide Through the Zoo of Biased SGD 21 Sep 2023 · 0 repositories
-
Approximate Heavy Tails in Offline (Multi-Pass) Stochastic Gradient Descent 21 Sep 2023 · 1 repository
-
Augmented Memory Replay-based Continual Learning Approaches for Network Intrusion Detection 21 Sep 2023 · 0 repositories
-
Automatic Clipping: Differentially Private Deep Learning Made Easier and Stronger 21 Sep 2023 · 1 repository
-
BayesTune: Bayesian Sparse Deep Model Fine-tuning 21 Sep 2023 · 1 repository
-
Beyond Exponential Graph: Communication-Efficient Topologies for Decentralized Learning via Finite-time Convergence 21 Sep 2023 · 1 repository
-
Differentially Private Image Classification by Learning Priors from Random Processes 21 Sep 2023 · 1 repository
-
Easy Bayesian Transfer Learning with Informative Priors 21 Sep 2023 · 0 repositories
-
Federated Learning with Client Subsampling, Data Heterogeneity, and Unbounded Smoothness: A New Algorithm and Lower Bounds 21 Sep 2023 · 1 repository
-
Global Convergence Analysis of Local SGD for Two-layer Neural Network without Overparameterization 21 Sep 2023 · 0 repositories
-
Implicit Bias of (Stochastic) Gradient Descent for Rank-1 Linear Neural Network 21 Sep 2023 · 0 repositories
-
KAKURENBO: Adaptively Hiding Samples in Deep Neural Network Training 21 Sep 2023 · 1 repository
-
Mean-field Langevin dynamics: Time-space discretization, stochastic gradient, and variance reduction 21 Sep 2023 · 0 repositories
-
(S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of Stability 21 Sep 2023 · 0 repositories
-
Preconditioned Federated Learning 20 Sep 2023 · 0 repositories · arXiv:2309.11378
-
On the different regimes of Stochastic Gradient Descent 19 Sep 2023 · 1 repository · arXiv:2309.10688Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 8 unverified (of 17 harvested samples)
-
Global Convergence of SGD For Logistic Loss on Two Layer Neural Nets 17 Sep 2023 · 0 repositories · arXiv:2309.09258
-
Stochastic Gradient Descent-like relaxation is equivalent to Metropolis dynamics in discrete optimization and inference problems 11 Sep 2023 · 1 repository · arXiv:2309.05337
-
Stochastic Gradient Descent outperforms Gradient Descent in recovering a high-dimensional signal in a glassy energy landscape 9 Sep 2023 · 0 repositories · arXiv:2309.04788
-
Convergence Analysis of Decentralized ASGD 7 Sep 2023 · 0 repositories · arXiv:2309.03754
-
AdaPlus: Integrating Nesterov Momentum and Precise Stepsize Adjustment on AdamW Basis 5 Sep 2023 · 1 repository · arXiv:2309.01966
-
Asymmetric Momentum: A Rethinking of Gradient Descent 5 Sep 2023 · 1 repository · arXiv:2309.02130
-
Corgi^2: A Hybrid Offline-Online Approach To Storage-Aware Data Shuffling For SGD 4 Sep 2023 · 0 repositories · arXiv:2309.01640
-
ABS-SGD: A Delayed Synchronous Stochastic Gradient Descent Algorithm with Adaptive Batch Size for Heterogeneous GPU Clusters 29 Aug 2023 · 0 repositories · arXiv:2308.15164
-
Bias-Aware Minimisation: Understanding and Mitigating Estimator Bias in Private SGD 23 Aug 2023 · 0 repositories · arXiv:2308.12018
-
Extended Linear Regression: A Kalman Filter Approach for Minimizing Loss via Area Under the Curve 23 Aug 2023 · 0 repositories · arXiv:2308.12280
-
When MiniBatch SGD Meets SplitFed Learning:Convergence Analysis and Performance Evaluation 23 Aug 2023 · 0 repositories · arXiv:2308.11953
-
Towards Understanding the Generalizability of Delayed Stochastic Gradient Descent 18 Aug 2023 · 0 repositories · arXiv:2308.09430
-
Hitting the High-Dimensional Notes: An ODE for SGD learning dynamics on GLMs and multi-index models 17 Aug 2023 · 0 repositories · arXiv:2308.08977
-
Max-affine regression via first-order methods 15 Aug 2023 · 0 repositories · arXiv:2308.08070
-
Law of Balance and Stationary Distribution of Stochastic Gradient Descent 13 Aug 2023 · 0 repositories · arXiv:2308.06671
-
Understanding the robustness difference between stochastic gradient descent and adaptive gradient methods 13 Aug 2023 · 1 repository · arXiv:2308.06703Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Adaptive SGD with Polyak stepsize and Line-search: Robust Convergence and Variance Reduction 11 Aug 2023 · 0 repositories · arXiv:2308.06058
-
Real-Time Progressive Learning: Accumulate Knowledge from Control with Neural-Network-Based Selective Memory 8 Aug 2023 · 0 repositories · arXiv:2308.04223
-
The Effect of SGD Batch Size on Autoencoder Learning: Sparsity, Sharpness, and Feature Learning 6 Aug 2023 · 0 repositories · arXiv:2308.03215
-
Eva: A General Vectorized Approximation Framework for Second-order Optimization 4 Aug 2023 · 0 repositories · arXiv:2308.02123
-
Get the Best of Both Worlds: Improving Accuracy and Transferability by Grassmann Class Representation 3 Aug 2023 · 1 repository · arXiv:2308.01547
-
Online covariance estimation for stochastic gradient descent under Markovian sampling 3 Aug 2023 · 0 repositories · arXiv:2308.01481
-
Toward Quantum Machine Translation of Syntactically Distinct Languages 31 Jul 2023 · 0 repositories · arXiv:2307.16576
-
The Marginal Value of Momentum for Small Learning Rate SGD 27 Jul 2023 · 0 repositories · arXiv:2307.15196
-
Function Value Learning: Adaptive Learning Rates Based on the Polyak Stepsize and Function Splitting in ERM 26 Jul 2023 · 0 repositories · arXiv:2307.14528
-
Relationship between Batch Size and Number of Steps Needed for Nonconvex Optimization of Stochastic Gradient Descent using Armijo Line Search 25 Jul 2023 · 0 repositories · arXiv:2307.13831
-
Convergence of SGD for Training Neural Networks with Sliced Wasserstein Losses 21 Jul 2023 · 0 repositories · arXiv:2307.11714
-
Stochastic Subgradient Methods with Guaranteed Global Stability in Nonsmooth Nonconvex Optimization 19 Jul 2023 · 0 repositories · arXiv:2307.10053
-
Weighted Averaged Stochastic Gradient Descent: Asymptotic Normality and Optimality 13 Jul 2023 · 0 repositories · arXiv:2307.06915
-
Mini-Batch Optimization of Contrastive Loss 12 Jul 2023 · 1 repository · arXiv:2307.05906
-
FedYolo: Augmenting Federated Learning with Pretrained Transformers 10 Jul 2023 · 0 repositories · arXiv:2307.04905
-
Bidirectional Looking with A Novel Double Exponential Moving Average to Adaptive and Non-adaptive Momentum Optimizers 2 Jul 2023 · 1 repository · arXiv:2307.00631Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples)
-
Federated Ensemble YOLOv5 -- A Better Generalized Object Detection Algorithm 30 Jun 2023 · 0 repositories · arXiv:2306.17829
-
Nonconvex Stochastic Bregman Proximal Gradient Method with Application to Deep Learning 26 Jun 2023 · 0 repositories · arXiv:2306.14522
-
Empirical Risk Minimization with Shuffled SGD: A Primal-Dual Perspective and Improved Bounds 21 Jun 2023 · 0 repositories · arXiv:2306.12498
-
InRank: Incremental Low-Rank Learning 20 Jun 2023 · 1 repository · arXiv:2306.11250
-
Adaptive Federated Learning with Auto-Tuned Clients 19 Jun 2023 · 1 repository · arXiv:2306.11201
-
Adaptive Strategies in Non-convex Optimization 17 Jun 2023 · 0 repositories · arXiv:2306.10278
-
Gradient is All You Need? 16 Jun 2023 · 2 repositories · arXiv:2306.09778
-
Evaluation and Optimization of Gradient Compression for Distributed Deep Learning 15 Jun 2023 · 1 repository · arXiv:2306.08881
-
Span-Selective Linear Attention Transformers for Effective and Robust Schema-Guided Dialogue State Tracking 15 Jun 2023 · 0 repositories · arXiv:2306.09340
-
Stochastic Re-weighted Gradient Descent via Distributionally Robust Optimization 15 Jun 2023 · 0 repositories · arXiv:2306.09222
-
When and Why Momentum Accelerates SGD:An Empirical Study 15 Jun 2023 · 0 repositories · arXiv:2306.09000
-
Beyond Implicit Bias: The Insignificance of SGD Noise in Online Learning 14 Jun 2023 · 0 repositories · arXiv:2306.08590
-
Kalman Filter for Online Classification of Non-Stationary Data 14 Jun 2023 · 0 repositories · arXiv:2306.08448
-
Noise Stability Optimization for Finding Flat Minima: A Hessian-based Regularization Approach 14 Jun 2023 · 1 repository · arXiv:2306.08553Syntology official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 15 harvested samples)
-
Exact Mean Square Linear Stability Analysis for SGD 13 Jun 2023 · 0 repositories · arXiv:2306.07850
-
Implicit Compressibility of Overparametrized Neural Networks Trained with Heavy-Tailed SGD 13 Jun 2023 · 1 repository · arXiv:2306.08125
-
Convergence of mean-field Langevin dynamics: Time and space discretization, stochastic gradient, and variance reduction 12 Jun 2023 · 0 repositories · arXiv:2306.07221
-
Fast Diffusion Model 12 Jun 2023 · 1 repository · arXiv:2306.06991Syntology official (archive's flag): 14 ran · 14 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 3 honoured, 0 violated, 5 with no contract checked; 6 where Syntology's instrument failed) · 3 unverified (of 17 harvested samples) · 9 pointer-only (licence)
-
Gaussian Membership Inference Privacy 12 Jun 2023 · 1 repository · arXiv:2306.07273
-
Correlated Noise in Epoch-Based Stochastic Gradient Descent: Implications for Weight Variances 8 Jun 2023 · 0 repositories · arXiv:2306.05300