Methods › General › Stochastic Optimization › SGD › Papers, page 4
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 4 of 21: papers 301 to 400 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Improving Implicit Regularization of SGD with Preconditioning for Least Square Problems 13 Mar 2024 · 0 repositories · arXiv:2403.08585
-
Do Deep Neural Network Solutions Form a Star Domain? 12 Mar 2024 · 1 repository · arXiv:2403.07968Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
Efficient Language Model Architectures for Differentially Private Federated Learning 12 Mar 2024 · 0 repositories · arXiv:2403.08100
-
SGD with Partial Hessian for Deep Neural Networks Optimization 5 Mar 2024 · 1 repository · arXiv:2403.02681
-
Shuffling Momentum Gradient Algorithm for Convex Optimization 5 Mar 2024 · 0 repositories · arXiv:2403.03180
-
SOFIM: Stochastic Optimization Using Regularized Fisher Information Matrix 5 Mar 2024 · 0 repositories · arXiv:2403.02833
-
Privacy of SGD under Gaussian or Heavy-Tailed Noise: Guarantees without Gradient Clipping 4 Mar 2024 · 0 repositories · arXiv:2403.02051
-
The Implicit Bias of Heterogeneity towards Invariance: A Study of Multi-Environment Matrix Sensing 3 Mar 2024 · 0 repositories · arXiv:2403.01420
-
Beyond Single-Model Views for Deep Learning: Optimization versus Generalizability of Stochastic Optimization Algorithms 1 Mar 2024 · 0 repositories · arXiv:2403.00574
-
Why Transformers Need Adam: A Hessian Perspective 26 Feb 2024 · 2 repositories · arXiv:2402.16788Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Effective Gradient Sample Size via Variation Estimation for Accelerating Sharpness aware Minimization 24 Feb 2024 · 0 repositories · arXiv:2403.08821
-
Dynamic Memory Based Adaptive Optimization 23 Feb 2024 · 0 repositories · arXiv:2402.15262
-
Iteration and Stochastic First-order Oracle Complexities of Stochastic Gradient Descent using Constant and Decaying Learning Rates 23 Feb 2024 · 1 repository · arXiv:2402.15344
-
SGD with Clipping is Secretly Estimating the Median Gradient 20 Feb 2024 · 0 repositories · arXiv:2402.12828
-
Training Artificial Neural Networks by Coordinate Search Algorithm 20 Feb 2024 · 0 repositories · arXiv:2402.12646
-
Adaptive Skeleton Graph Decoding 19 Feb 2024 · 0 repositories · arXiv:2402.12280
-
Communication-Efficient Distributed Learning with Local Immediate Error Compensation 19 Feb 2024 · 0 repositories · arXiv:2402.11857
-
Diagonalisation SGD: Fast & Convergent SGD for Non-Differentiable Models via Reparameterisation and Smoothing 19 Feb 2024 · 1 repository · arXiv:2402.11752
-
OptEx: Expediting First-Order Optimization with Approximately Parallelized Iterations 18 Feb 2024 · 1 repository · arXiv:2402.11427Syntology official (archive's flag): 7 ran · 7 ran (of which 1 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 5 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Revisiting Zeroth-Order Optimization for Memory-Efficient LLM Fine-Tuning: A Benchmark 18 Feb 2024 · 1 repository · arXiv:2402.11592Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 6 where Syntology's instrument failed) · 4 unverified (of 13 harvested samples) · 13 pointer-only (licence)
-
Implicit Bias in Noisy-SGD: With Applications to Differentially Private Training 13 Feb 2024 · 0 repositories · arXiv:2402.08344
-
Differentially Private Zeroth-Order Methods for Scalable Large Language Model Finetuning 12 Feb 2024 · 0 repositories · arXiv:2402.07818
-
Tuning-Free Stochastic Optimization 12 Feb 2024 · 0 repositories · arXiv:2402.07793
-
Parameter Symmetry and Noise Equilibrium of Stochastic Gradient Descent 11 Feb 2024 · 0 repositories · arXiv:2402.07193
-
Should I try multiple optimizers when fine-tuning pre-trained Transformers for NLP tasks? Should I tune their hyperparameters? 10 Feb 2024 · 0 repositories · arXiv:2402.06948
-
Exploring Learning Complexity for Efficient Downstream Dataset Pruning 8 Feb 2024 · 0 repositories · arXiv:2402.05356
-
On the Convergence of Zeroth-Order Federated Tuning for Large Language Models 8 Feb 2024 · 1 repository · arXiv:2402.05926
-
AdaBatchGrad: Combining Adaptive Batch Size and Adaptive Step Size 7 Feb 2024 · 0 repositories · arXiv:2402.05264
-
Curvature-Informed SGD via General Purpose Lie-Group Preconditioners 7 Feb 2024 · 1 repository · arXiv:2402.04553
-
Feature learning as alignment: a structural property of gradient descent in non-linear neural networks 7 Feb 2024 · 0 repositories · arXiv:2402.05271
-
Learning Operators with Stochastic Gradient Descent in General Hilbert Spaces 7 Feb 2024 · 0 repositories · arXiv:2402.04691
-
Non-convergence to global minimizers for Adam and stochastic gradient descent optimization and constructions of local minimizers in the training of artificial neural networks 7 Feb 2024 · 0 repositories · arXiv:2402.05155
-
Shadowheart SGD: Distributed Asynchronous SGD with Optimal Time Complexity Under Arbitrary Computation and Communication Heterogeneity 7 Feb 2024 · 0 repositories · arXiv:2402.04785
-
Can We Remove the Square-Root in Adaptive Gradient Methods? A Second-Order Perspective 5 Feb 2024 · 2 repositories · arXiv:2402.03496Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 2 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 9 pointer-only (licence)
-
Non-asymptotic Analysis of Biased Adaptive Stochastic Approximation 5 Feb 2024 · 1 repository · arXiv:2402.02857
-
Vanishing Feature: Diagnosing Model Merging and Beyond 5 Feb 2024 · 2 repositories · arXiv:2402.05966Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Riemannian Preconditioned LoRA for Fine-Tuning Foundation Models 4 Feb 2024 · 1 repository · arXiv:2402.02347Syntology official (archive's flag): 12 ran · 15 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 6 unverified (of 21 harvested samples) · 3 pointer-only (licence)
-
Momentum Does Not Reduce Stochastic Noise in Stochastic Gradient Descent 4 Feb 2024 · 0 repositories · arXiv:2402.02325
-
Emergence of heavy tails in homogenized stochastic gradient descent 2 Feb 2024 · 0 repositories · arXiv:2402.01382
-
Enhancing Stochastic Gradient Descent: A Unified Framework and Novel Acceleration Methods for Faster Convergence 2 Feb 2024 · 0 repositories · arXiv:2402.01515
-
Truncated Non-Uniform Quantization for Distributed SGD 2 Feb 2024 · 0 repositories · arXiv:2402.01160
-
Understanding Adam Optimizer via Online Learning of Updates: Adam is FTRL in Disguise 2 Feb 2024 · 0 repositories · arXiv:2402.01567
-
Comparing Spectral Bias and Robustness For Two-Layer Neural Networks: SGD vs Adaptive Random Fourier Features 1 Feb 2024 · 0 repositories · arXiv:2402.00332
-
On the O((√(d))/(T^(1/4))) Convergence Rate of RMSProp and Its Momentum Extension Measured by ℓ₁ Norm 1 Feb 2024 · 0 repositories · arXiv:2402.00389
-
HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy 26 Jan 2024 · 1 repository · arXiv:2401.15207Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
On Principled Local Optimization Methods for Federated Learning 24 Jan 2024 · 0 repositories · arXiv:2401.13216
-
A Precise Characterization of SGD Stability Using Loss Surface Geometry 22 Jan 2024 · 0 repositories · arXiv:2401.12332
-
Momentum-SAM: Sharpness Aware Minimization without Computational Overhead 22 Jan 2024 · 1 repository · arXiv:2401.12033Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
The Dimension Strikes Back with Gradients: Generalization of Gradient Methods in Stochastic Convex Optimization 22 Jan 2024 · 0 repositories · arXiv:2401.12058
-
Understanding the Generalization Benefits of Late Learning Rate Decay 21 Jan 2024 · 0 repositories · arXiv:2401.11600
-
Communication Efficient and Provable Federated Unlearning 19 Jan 2024 · 1 repository · arXiv:2401.11018
-
Asynchronous Local-SGD Training for Language Modeling 17 Jan 2024 · 1 repository · arXiv:2401.09135
-
Central Limit Theorem for Two-Timescale Stochastic Approximation with Markovian Noise: Theory and Applications 17 Jan 2024 · 0 repositories · arXiv:2401.09339
-
Stabilizing Sharpness-aware Minimization Through A Simple Renormalization Strategy 14 Jan 2024 · 2 repositories · arXiv:2401.07250
-
An ADRC-Incorporated Stochastic Gradient Descent Algorithm for Latent Factor Analysis 13 Jan 2024 · 0 repositories · arXiv:2401.07012
-
(Accelerated) Noise-adaptive Stochastic Heavy-Ball Momentum 12 Jan 2024 · 1 repository · arXiv:2401.06738
-
Correlated Quantization for Faster Nonconvex Distributed Optimization 10 Jan 2024 · 0 repositories · arXiv:2401.05518
-
AA-DLADMM: An Accelerated ADMM-based Framework for Training Deep Neural Networks 8 Jan 2024 · 0 repositories · arXiv:2401.03619
-
On the numerical reliability of nonsmooth autodiff: a MaxPool case study 5 Jan 2024 · 1 repository · arXiv:2401.02736
-
Ravnest: Decentralized Asynchronous Training on Heterogeneous Devices 3 Jan 2024 · 0 repositories · arXiv:2401.01728
-
Online Tensor Inference 28 Dec 2023 · 0 repositories · arXiv:2312.17111
-
On the Trajectories of SGD Without Replacement 26 Dec 2023 · 0 repositories · arXiv:2312.16143
-
ZO-AdaMU Optimizer: Adapting Perturbation by the Momentum and Uncertainty in Zeroth-order Optimization 23 Dec 2023 · 1 repository · arXiv:2312.15184Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Accelerated Convergence of Stochastic Heavy Ball Method under Anisotropic Gradient Noise 22 Dec 2023 · 0 repositories · arXiv:2312.14567
-
On the Convergence of Loss and Uncertainty-based Active Learning Algorithms 21 Dec 2023 · 1 repository · arXiv:2312.13927
-
Parallel Trust-Region Approaches in Neural Network Training: Beyond Traditional Methods 21 Dec 2023 · 0 repositories · arXiv:2312.13677
-
Contractive error feedback for gradient compression 13 Dec 2023 · 0 repositories · arXiv:2312.08538
-
Revisiting the Last-Iterate Convergence of Stochastic Gradient Methods 13 Dec 2023 · 0 repositories · arXiv:2312.08531
-
TOD-Flow: Modeling the Structure of Task-Oriented Dialogues 7 Dec 2023 · 1 repository · arXiv:2312.04668Syntology official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 10 harvested samples) · 10 pointer-only (licence)
-
Convergence Rates for Stochastic Approximation: Biased Noise with Unbounded Variance, and Applications 5 Dec 2023 · 0 repositories · arXiv:2312.02828
-
AGD: an Auto-switchable Optimizer using Stepwise Gradient Difference for Preconditioning Matrix 4 Dec 2023 · 2 repositories · arXiv:2312.01658Syntology official (archive's flag): 1 ran · 2 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 1 pointer-only (licence)
-
Unlocking optimal batch size schedules using continuous-time control and perturbation theory 4 Dec 2023 · 0 repositories · arXiv:2312.01898
-
Can We Learn Communication-Efficient Optimizers? 2 Dec 2023 · 1 repository · arXiv:2312.02204
-
Temperature Balancing, Layer-wise Weight Analysis, and Neural Network Training 1 Dec 2023 · 1 repository · arXiv:2312.00359Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
The Feature Speed Formula: a flexible approach to scale hyper-parameters of deep neural networks 30 Nov 2023 · 0 repositories · arXiv:2311.18718
-
Critical Influence of Overparameterization on Sharpness-aware Minimization 29 Nov 2023 · 1 repository · arXiv:2311.17539Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 9 harvested samples) · 9 pointer-only (licence)
-
In Search of a Data Transformation That Accelerates Neural Field Training 28 Nov 2023 · 1 repository · arXiv:2311.17094
-
MAST: Model-Agnostic Sparsified Training 27 Nov 2023 · 1 repository · arXiv:2311.16086Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Scheduling and Communication Schemes for Decentralized Federated Learning 27 Nov 2023 · 0 repositories · arXiv:2311.16021
-
Online Importance Sampling for Stochastic Gradient Optimization 24 Nov 2023 · 0 repositories · arXiv:2311.14468
-
Risk Bounds of Accelerated SGD for Overparameterized Linear Regression 23 Nov 2023 · 0 repositories · arXiv:2311.14222
-
Weight fluctuations in (deep) linear neural networks and a derivation of the inverse-variance flatness relation 23 Nov 2023 · 0 repositories · arXiv:2311.14120
-
Sample as You Infer: Predictive Coding With Langevin Dynamics 22 Nov 2023 · 0 repositories · arXiv:2311.13664
-
Are Normalizing Flows the Key to Unlocking the Exponential Mechanism? 15 Nov 2023 · 1 repository · arXiv:2311.09200
-
Data-Aware Gradient Compression for FL in Communication-Constrained Mobile Computing 13 Nov 2023 · 0 repositories · arXiv:2311.07324
-
Compressed and Sparse Models for Non-Convex Decentralized Learning 9 Nov 2023 · 0 repositories · arXiv:2311.05760
-
Outliers with Opposing Signals Have an Outsized Effect on Neural Network Optimization 7 Nov 2023 · 0 repositories · arXiv:2311.04163
-
Signal Processing Meets SGD: From Momentum to Filter 6 Nov 2023 · 1 repository · arXiv:2311.02818
-
Asynchronous SGD on Graphs: a Unified Framework for Asynchronous Decentralized and Federated Optimization 1 Nov 2023 · 0 repositories · arXiv:2311.00465
-
AsGrad: A Sharp Unified Analysis of Asynchronous-SGD Algorithms 31 Oct 2023 · 0 repositories · arXiv:2310.20452
-
Information-Theoretic Trust Regions for Stochastic Gradient-Based Optimization 31 Oct 2023 · 1 repository · arXiv:2310.20574
-
Escaping Saddle Points in Heterogeneous Federated Learning via Distributed SGD with Communication Compression 29 Oct 2023 · 0 repositories · arXiv:2310.19059
-
High-probability Convergence Bounds for Nonlinear Stochastic Gradient Descent Under Heavy-tailed Noise 28 Oct 2023 · 0 repositories · arXiv:2310.18784
-
Linear Mode Connectivity in Sparse Neural Networks 28 Oct 2023 · 0 repositories · arXiv:2310.18769
-
Benign Oscillation of Stochastic Gradient Descent with Large Learning Rates 26 Oct 2023 · 0 repositories · arXiv:2310.17074
-
Grokking Beyond Neural Networks: An Empirical Exploration with Model Complexity 26 Oct 2023 · 1 repository · arXiv:2310.17247
-
A model for multi-attack classification to improve intrusion detection performance using deep learning approaches 25 Oct 2023 · 0 repositories · arXiv:2310.16380
-
Probabilistic Integral Circuits 25 Oct 2023 · 0 repositories · arXiv:2310.16986
-
Learning From Free-Text Human Feedback -- Collect New Datasets Or Extend Existing Ones? 24 Oct 2023 · 1 repository · arXiv:2310.15758Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Studying K-FAC Heuristics by Viewing Adam through a Second-Order Lens 23 Oct 2023 · 1 repository · arXiv:2310.14963Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 9 unverified (of 16 harvested samples) · 16 pointer-only (licence)