Methods › General › Stochastic Optimization › SGD › Papers, page 7
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 7 of 21: papers 601 to 700 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
The Monge Gap: A Regularizer to Learn All Transport Maps 9 Feb 2023 · 0 repositories · arXiv:2302.04953
-
DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size Schedule 8 Feb 2023 · 1 repository · arXiv:2302.12022Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
Easy Learning from Label Proportions 6 Feb 2023 · 1 repository · arXiv:2302.03115
-
On the Convergence of Federated Averaging with Cyclic Client Participation 6 Feb 2023 · 0 repositories · arXiv:2302.03109
-
Toward Large Kernel Models 6 Feb 2023 · 1 repository · arXiv:2302.02605
-
z-SignFedAvg: A Unified Stochastic Sign-based Compression for Federated Learning 6 Feb 2023 · 0 repositories · arXiv:2302.02589
-
Quantized Distributed Training of Large Models with Convergence Guarantees 5 Feb 2023 · 0 repositories · arXiv:2302.02390
-
Selecting the Best Optimizers for Deep Learning based Medical Image Segmentation 5 Feb 2023 · 0 repositories · arXiv:2302.02289
-
On Suppressing Range of Adaptive Stepsizes of Adam to Improve Generalisation Performance 2 Feb 2023 · 0 repositories · arXiv:2302.01029
-
Coordinating Distributed Example Orders for Provably Accelerated Training 2 Feb 2023 · 1 repository · arXiv:2302.00845Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
STEP: Learning N:M Structured Sparsity Masks from Scratch with Precondition 2 Feb 2023 · 1 repository · arXiv:2302.01172
-
Dissecting the Effects of SGD Noise in Distinct Regimes of Deep Learning 31 Jan 2023 · 0 repositories · arXiv:2301.13703
-
Schema-Guided Semantic Accuracy: Faithfulness in Task-Oriented Dialogue Response Generation 29 Jan 2023 · 1 repository · arXiv:2301.12568
-
CyclicFL: A Cyclic Model Pre-Training Approach to Efficient Federated Learning 28 Jan 2023 · 0 repositories · arXiv:2301.12193
-
On the Lipschitz Constant of Deep Networks and Double Descent 28 Jan 2023 · 1 repository · arXiv:2301.12309
-
Algorithmic Stability of Heavy-Tailed SGD with General Loss Functions 27 Jan 2023 · 0 repositories · arXiv:2301.11885
-
Trainable Activations for Image Classification 26 Jan 2023 · 1 repository
-
Transfer Learning in Deep Learning Models for Building Load Forecasting: Case of Limited Data 25 Jan 2023 · 2 repositories · arXiv:2301.10663
-
Read the Signs: Towards Invariance to Gradient Descent's Hyperparameter Initialization 24 Jan 2023 · 1 repository · arXiv:2301.10133
-
Genetically Modified Wolf Optimization with Stochastic Gradient Descent for Optimising Deep Neural Networks 21 Jan 2023 · 0 repositories · arXiv:2301.08950
-
ScaDLES: Scalable Deep Learning over Streaming data at the Edge 21 Jan 2023 · 1 repository · arXiv:2301.08897
-
Learning-Rate-Free Learning by D-Adaptation 18 Jan 2023 · 1 repository · arXiv:2301.07733Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Convergence of First-Order Algorithms for Meta-Learning with Moreau Envelopes 17 Jan 2023 · 0 repositories · arXiv:2301.06806
-
Expected Gradients of Maxout Networks and Consequences to Parameter Initialization 17 Jan 2023 · 1 repository · arXiv:2301.06956
-
CEDAS: A Compressed Decentralized Stochastic Gradient Method with Improved Convergence 14 Jan 2023 · 0 repositories · arXiv:2301.05872
-
Toward Theoretical Guidance for Two Common Questions in Practical Cross-Validation based Hyperparameter Selection 12 Jan 2023 · 0 repositories · arXiv:2301.05131
-
Training trajectories, mini-batch losses and the curious role of the learning rate 5 Jan 2023 · 0 repositories · arXiv:2301.02312
-
On the Convergence of Stochastic Gradient Descent in Low-precision Number Formats 4 Jan 2023 · 0 repositories · arXiv:2301.01651
-
Temporal Difference Learning with Compressed Updates: Error-Feedback meets Reinforcement Learning 3 Jan 2023 · 0 repositories · arXiv:2301.00944
-
GPU accelerated matrix factorization of large scale data using block based approach 2 Jan 2023 · 0 repositories · arXiv:2304.13724
-
Online Statistical Inference for Contextual Bandits via Stochastic Gradient Descent 30 Dec 2022 · 0 repositories · arXiv:2212.14883
-
Visualizing Information Bottleneck through Variational Inference 24 Dec 2022 · 0 repositories · arXiv:2212.12667
-
When Do Curricula Work in Federated Learning? 24 Dec 2022 · 1 repository · arXiv:2212.12712
-
AnyTOD: A Programmable Task-Oriented Dialog System 20 Dec 2022 · 0 repositories · arXiv:2212.09939
-
Improving Levenberg-Marquardt Algorithm for Neural Networks 17 Dec 2022 · 1 repository · arXiv:2212.08769
-
Huber-energy measure quantization 15 Dec 2022 · 0 repositories · arXiv:2212.08162
-
Neuroevolution of Physics-Informed Neural Nets: Benchmark Problems and Comparative Results 15 Dec 2022 · 0 repositories · arXiv:2212.07624
-
Generalizing DP-SGD with Shuffling and Batch Clipping 12 Dec 2022 · 0 repositories · arXiv:2212.05796
-
Covariance Estimators for the ROOT-SGD Algorithm in Online Learning 2 Dec 2022 · 0 repositories · arXiv:2212.01259
-
Investigating certain choices of CNN configurations for brain lesion segmentation 2 Dec 2022 · 0 repositories · arXiv:2212.01235
-
Disentangling the Mechanisms Behind Implicit Regularization in SGD 29 Nov 2022 · 1 repository · arXiv:2211.15853
-
Stochastic Steffensen method 28 Nov 2022 · 0 repositories · arXiv:2211.15310
-
Learning Compact Features via In-Training Representation Alignment 23 Nov 2022 · 0 repositories · arXiv:2211.13332
-
Mutual Information Learned Regressor: an Information-theoretic Viewpoint of Training Regression Systems 23 Nov 2022 · 0 repositories · arXiv:2211.12685
-
ModelDiff: A Framework for Comparing Learning Algorithms 22 Nov 2022 · 1 repository · arXiv:2211.12491Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 8 harvested samples)
-
Scaling Up Dataset Distillation to ImageNet-1K with Constant Memory 19 Nov 2022 · 2 repositories · arXiv:2211.10586Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Two Facets of SDE Under an Information-Theoretic Lens: Generalization of SGD via Training Trajectories and via Terminal States 19 Nov 2022 · 0 repositories · arXiv:2211.10691
-
How to Fine-Tune Vision Models with SGD 17 Nov 2022 · 0 repositories · arXiv:2211.09359
-
SketchySGD: Reliable Stochastic Optimization via Randomized Curvature Estimates 16 Nov 2022 · 1 repository · arXiv:2211.08597
-
Empirical Study on Optimizer Selection for Out-of-Distribution Generalization 15 Nov 2022 · 1 repository · arXiv:2211.08583Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples) · 1 pointer-only (licence)
-
REPAIR: REnormalizing Permuted Activations for Interpolation Repair 15 Nov 2022 · 1 repository · arXiv:2211.08403
-
Selective Memory Recursive Least Squares: Recast Forgetting into Memory in RBF Neural Network Based Real-Time Learning 15 Nov 2022 · 0 repositories · arXiv:2211.07909
-
Alternating Implicit Projected SGD and Its Efficient Variants for Equality-constrained Bilevel Optimization 14 Nov 2022 · 1 repository · arXiv:2211.07096
-
Multi-Epoch Matrix Factorization Mechanisms for Private Machine Learning 12 Nov 2022 · 1 repository · arXiv:2211.06530
-
Variants of SGD for Lipschitz Continuous Loss Functions in Low-Precision Environments 9 Nov 2022 · 0 repositories · arXiv:2211.04655
-
AskewSGD : An Annealed interval-constrained Optimisation method to train Quantized Neural Networks 7 Nov 2022 · 0 repositories · arXiv:2211.03741
-
Accelerating Parallel Stochastic Gradient Descent via Non-blocking Mini-batches 2 Nov 2022 · 0 repositories · arXiv:2211.00889
-
Large deviations rates for stochastic gradient descent with strongly convex functions 2 Nov 2022 · 0 repositories · arXiv:2211.00969
-
RCD-SGD: Resource-Constrained Distributed SGD in Heterogeneous Environment via Submodular Partitioning 2 Nov 2022 · 0 repositories · arXiv:2211.00839
-
Strong Lottery Ticket Hypothesis with ε--perturbation 29 Oct 2022 · 0 repositories · arXiv:2210.16589
-
Flatter, faster: scaling momentum for optimal speedup of SGD 28 Oct 2022 · 0 repositories · arXiv:2210.16400
-
Stochastic Mirror Descent in Average Ensemble Models 27 Oct 2022 · 0 repositories · arXiv:2210.15323
-
Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models 25 Oct 2022 · 0 repositories · arXiv:2210.14199
-
Adaptive Top-K in SGD for Communication-Efficient Distributed Learning 24 Oct 2022 · 0 repositories · arXiv:2210.13532
-
Decentralized Stochastic Bilevel Optimization with Improved per-Iteration Complexity 23 Oct 2022 · 0 repositories · arXiv:2210.12839
-
K-SAM: Sharpness-Aware Minimization at the Speed of SGD 23 Oct 2022 · 0 repositories · arXiv:2210.12864
-
When Expressivity Meets Trainability: Fewer than n Neurons Can Work 21 Oct 2022 · 0 repositories · arXiv:2210.12001
-
Global Convergence of SGD On Two Layer Neural Nets 20 Oct 2022 · 0 repositories · arXiv:2210.11452
-
Large-batch Optimization for Dense Visual Predictions 20 Oct 2022 · 1 repository · arXiv:2210.11078
-
Local SGD in Overparameterized Linear Regression 20 Oct 2022 · 0 repositories · arXiv:2210.11562
-
DPIS: An Enhanced Mechanism for Differentially Private SGD with Importance Sampling 18 Oct 2022 · 1 repository · arXiv:2210.09634
-
Linear Scalarization for Byzantine-robust learning on non-IID data 15 Oct 2022 · 0 repositories · arXiv:2210.08287
-
Communication-Efficient Topologies for Decentralized Learning with O(1) Consensus Rate 14 Oct 2022 · 1 repository · arXiv:2210.07881
-
When Adversarial Training Meets Vision Transformers: Recipes from Training to Architecture 14 Oct 2022 · 1 repository · arXiv:2210.07540
-
From Gradient Flow on Population Loss to Learning with Stochastic Gradient Descent 13 Oct 2022 · 0 repositories · arXiv:2210.06705
-
Mean-field analysis for heavy ball methods: Dropout-stability, connectivity, and global convergence 13 Oct 2022 · 0 repositories · arXiv:2210.06819
-
Wasserstein Barycenter-based Model Fusion and Linear Mode Connectivity of Neural Networks 13 Oct 2022 · 1 repository · arXiv:2210.06671
-
AdaNorm: Adaptive Gradient Norm Correction based Optimizer for CNNs 12 Oct 2022 · 1 repository · arXiv:2210.06364
-
Improving information retention in large scale online continual learning 12 Oct 2022 · 0 repositories · arXiv:2210.06401
-
Rigorous dynamical mean field theory for stochastic gradient descent methods 12 Oct 2022 · 1 repository · arXiv:2210.06591
-
A Kernel-Based View of Language Model Fine-Tuning 11 Oct 2022 · 1 repository · arXiv:2210.05643Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples) · 1 pointer-only (licence)
-
SGD with Large Step Sizes Learns Sparse Features 11 Oct 2022 · 1 repository · arXiv:2210.05337
-
Fast Hierarchical Learning for Few-Shot Object Detection 10 Oct 2022 · 0 repositories · arXiv:2210.05008
-
Dissecting adaptive methods in GANs 9 Oct 2022 · 0 repositories · arXiv:2210.04319
-
STSyn: Speeding Up Local SGD with Straggler-Tolerant Synchronization 6 Oct 2022 · 0 repositories · arXiv:2210.03521
-
Scaling up Stochastic Gradient Descent for Non-convex Optimisation 6 Oct 2022 · 0 repositories · arXiv:2210.02882
-
Unmasking the Lottery Ticket Hypothesis: What's Encoded in a Winning Ticket's Mask? 6 Oct 2022 · 0 repositories · arXiv:2210.03044
-
Learning an Invertible Output Mapping Can Mitigate Simplicity Bias in Neural Networks 4 Oct 2022 · 0 repositories · arXiv:2210.01360
-
Distributed Non-Convex Optimization with One-Bit Compressors on Heterogeneous Data: Efficient and Resilient Algorithms 3 Oct 2022 · 0 repositories · arXiv:2210.00665
-
Downlink Compression Improves TopK Sparsification 30 Sep 2022 · 0 repositories · arXiv:2209.15203
-
Momentum Tracking: Momentum Acceleration for Decentralized Deep Learning on Heterogeneous Data 30 Sep 2022 · 0 repositories · arXiv:2209.15505
-
NAG-GS: Semi-Implicit, Accelerated and Robust Stochastic Optimizer 29 Sep 2022 · 2 repositories · arXiv:2209.14937Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
Neural Networks Efficiently Learn Low-Dimensional Representations with SGD 29 Sep 2022 · 0 repositories · arXiv:2209.14863
-
Statistical Learning and Inverse Problems: A Stochastic Gradient Approach 29 Sep 2022 · 0 repositories · arXiv:2209.14967
-
On the Stability Analysis of Open Federated Learning Systems 25 Sep 2022 · 0 repositories · arXiv:2209.12307
-
Error Mitigation-Aided Optimization of Parameterized Quantum Circuits: Convergence Analysis 23 Sep 2022 · 0 repositories · arXiv:2209.11514
-
Robust Collaborative Learning with Linear Gradient Overhead 22 Sep 2022 · 1 repository · arXiv:2209.10931Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples)
-
Generalization Bounds for Stochastic Gradient Descent via Localized ε-Covers 19 Sep 2022 · 0 repositories · arXiv:2209.08951
-
Stability and Generalization Analysis of Gradient Methods for Shallow Neural Networks 19 Sep 2022 · 0 repositories · arXiv:2209.09298
-
Empirical Analysis on Top-k Gradient Sparsification for Distributed Deep Learning in a Supercomputing Environment 18 Sep 2022 · 0 repositories · arXiv:2209.08497