Methods › General › Stochastic Optimization › SGD › Papers, page 15
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 15 of 21: papers 1,401 to 1,500 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Understanding the Difficulty of Training Transformers 17 Apr 2020 · 2 repositories · arXiv:2004.08249Syntology official (archive's flag): 1 ran · 4 ran (of which 2 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples)
-
On Learning Rates and Schrödinger Operators 15 Apr 2020 · 0 repositories · arXiv:2004.06977
-
Detached Error Feedback for Distributed SGD with Random Sparsification 11 Apr 2020 · 0 repositories · arXiv:2004.05298
-
FDA: Fourier Domain Adaptation for Semantic Segmentation 11 Apr 2020 · 3 repositories · arXiv:2004.05498Syntology official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Convergence rates and approximation results for SGD and its continuous-time counterpart 8 Apr 2020 · 0 repositories · arXiv:2004.04193
-
Stopping Criteria for, and Strong Convergence of, Stochastic Gradient Descent on Bottou-Curtis-Nocedal Functions 1 Apr 2020 · 0 repositories · arXiv:2004.00475
-
SS-IL: Separated Softmax for Incremental Learning 31 Mar 2020 · 0 repositories · arXiv:2003.13947
-
Stochastic Proximal Gradient Algorithm with Minibatches. Application to Large Scale Learning Models 30 Mar 2020 · 0 repositories · arXiv:2003.13332
-
A Hybrid-Order Distributed SGD Method for Non-Convex Optimization to Balance Communication Overhead, Computational Complexity, and Convergence Rate 27 Mar 2020 · 0 repositories · arXiv:2003.12423
-
On Infinite-Width Hypernetworks 27 Mar 2020 · 1 repository · arXiv:2003.12193
-
Convergence of Recursive Stochastic Algorithms using Wasserstein Divergence 25 Mar 2020 · 0 repositories · arXiv:2003.11403
-
Understanding the Effects of Data Parallelism and Sparsity on Neural Network Training 25 Mar 2020 · 0 repositories · arXiv:2003.11316Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Pipelined Backpropagation at Scale: Training Large Models without Batches 25 Mar 2020 · 0 repositories · arXiv:2003.11666
-
FedSel: Federated SGD under Local Differential Privacy with Top-k Dimension Selection 24 Mar 2020 · 0 repositories · arXiv:2003.10637
-
Finite-Time Analysis of Stochastic Gradient Descent under Markov Randomness 24 Mar 2020 · 0 repositories · arXiv:2003.10973
-
Online stochastic gradient descent on non-convex losses from high-dimensional inference 23 Mar 2020 · 0 repositories · arXiv:2003.10409
-
A Unified Theory of Decentralized SGD with Changing Topology and Local Updates 23 Mar 2020 · 0 repositories · arXiv:2003.10422
-
Slow and Stale Gradients Can Win the Race 23 Mar 2020 · 0 repositories · arXiv:2003.10579
-
NeCPD: An Online Tensor Decomposition with Optimal Stochastic Gradient Descent 18 Mar 2020 · 0 repositories · arXiv:2003.08844
-
Weak and Strong Gradient Directions: Explaining Memorization, Generalization, and Hardness of Examples at Scale 16 Mar 2020 · 0 repositories · arXiv:2003.07422
-
Investigating Generalization in Neural Networks under Optimally Evolved Training Perturbations 14 Mar 2020 · 1 repository · arXiv:2003.06646
-
Can Implicit Bias Explain Generalization? Stochastic Convex Optimization as a Case Study 13 Mar 2020 · 0 repositories · arXiv:2003.06152
-
Machine Learning on Volatile Instances 12 Mar 2020 · 0 repositories · arXiv:2003.05649
-
A Mean-field Analysis of Deep ResNet and Beyond: Towards Provable Optimization Via Overparameterization From Depth 11 Mar 2020 · 0 repositories · arXiv:2003.05508
-
ShadowSync: Performing Synchronization in the Background for Highly Scalable Distributed Training 7 Mar 2020 · 0 repositories · arXiv:2003.03477
-
A Simple Convergence Proof of Adam and Adagrad 5 Mar 2020 · 0 repositories · arXiv:2003.02395
-
Neural Kernels Without Tangents 4 Mar 2020 · 2 repositories · arXiv:2003.02237Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Buffered Asynchronous SGD for Byzantine Learning 2 Mar 2020 · 0 repositories · arXiv:2003.00937
-
Iterative Averaging in the Quest for Best Test Error 2 Mar 2020 · 0 repositories · arXiv:2003.01247
-
On the Global Convergence of Training Deep Linear ResNets 2 Mar 2020 · 0 repositories · arXiv:2003.01094
-
Loss landscapes and optimization in over-parameterized non-linear systems and neural networks 29 Feb 2020 · 0 repositories · arXiv:2003.00307
-
Distributed Momentum for Byzantine-resilient Learning 28 Feb 2020 · 1 repository · arXiv:2003.00010Syntology official: harvested, nothing ran · 0 ran · 3 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Do optimization methods in deep learning applications matter? 28 Feb 2020 · 0 repositories · arXiv:2002.12642
-
Fast and Three-rious: Speeding Up Weak Supervision with Triplet Methods 27 Feb 2020 · 1 repository · arXiv:2002.11955
-
On Biased Compression for Distributed Learning 27 Feb 2020 · 0 repositories · arXiv:2002.12410
-
LASG: Lazily Aggregated Stochastic Gradients for Communication-Efficient Distributed Learning 26 Feb 2020 · 1 repository · arXiv:2002.11360
-
Moniqua: Modulo Quantized Communication in Decentralized SGD 26 Feb 2020 · 0 repositories · arXiv:2002.11787
-
Non-asymptotic bounds for stochastic optimization with biased noisy gradient oracles 26 Feb 2020 · 0 repositories · arXiv:2002.11440
-
Stagewise Enlargement of Batch Size for SGD-based Learning 26 Feb 2020 · 0 repositories · arXiv:2002.11601
-
Adaptive Distributed Stochastic Gradient Descent for Minimizing Delay in the Presence of Stragglers 25 Feb 2020 · 0 repositories · arXiv:2002.11005
-
Stochastic-Sign SGD for Federated Learning with Theoretical Guarantees 25 Feb 2020 · 0 repositories · arXiv:2002.10940
-
Closing the convergence gap of SGD without replacement 24 Feb 2020 · 0 repositories · arXiv:2002.10400
-
HYDRA: Pruning Adversarially Robust Neural Networks 24 Feb 2020 · 4 repositories · arXiv:2002.10509Syntology community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 15 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (of 19 harvested samples) · 17 pointer-only (licence)
-
Scheduled Restart Momentum for Accelerated Stochastic Gradient Descent 24 Feb 2020 · 1 repository · arXiv:2002.10583Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Stochastic Polyak Step-size for SGD: An Adaptive Learning Rate for Fast Convergence 24 Feb 2020 · 1 repository · arXiv:2002.10542
-
Improve SGD Training via Aligning Mini-batches 23 Feb 2020 · 0 repositories · arXiv:2002.09917
-
Learning to Continually Learn 21 Feb 2020 · 5 repositories · arXiv:2002.09571Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Overlap Local-SGD: An Algorithmic Approach to Hide Communication Delays in Distributed SGD 21 Feb 2020 · 1 repository · arXiv:2002.09539
-
The Break-Even Point on Optimization Trajectories of Deep Neural Networks 21 Feb 2020 · 0 repositories · arXiv:2002.09572
-
Bounding the expected run-time of nonconvex optimization with early stopping 20 Feb 2020 · 0 repositories · arXiv:2002.08856
-
Embedding Graph Auto-Encoder for Graph Clustering 20 Feb 2020 · 0 repositories · arXiv:2002.08643
-
Stochastic Runge-Kutta methods and adaptive SGD-G2 stochastic gradient descent 20 Feb 2020 · 0 repositories · arXiv:2002.09304
-
Distributed Optimization over Block-Cyclic Data 18 Feb 2020 · 0 repositories · arXiv:2002.07454
-
Is Local SGD Better than Minibatch SGD? 18 Feb 2020 · 0 repositories · arXiv:2002.07839
-
Fast Convergence for Langevin Diffusion with Manifold Structure 13 Feb 2020 · 0 repositories · arXiv:2002.05576
-
Learning Halfspaces with Massart Noise Under Structured Distributions 13 Feb 2020 · 0 repositories · arXiv:2002.05632
-
A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima 10 Feb 2020 · 0 repositories · arXiv:2002.03495
-
Online Covariance Matrix Estimation in Stochastic Gradient Descent 10 Feb 2020 · 0 repositories · arXiv:2002.03979
-
Federated Learning of a Mixture of Global and Local Models 10 Feb 2020 · 0 repositories · arXiv:2002.05516
-
Semi-Implicit Back Propagation 10 Feb 2020 · 0 repositories · arXiv:2002.03516
-
Better Theory for SGD in the Nonconvex World 9 Feb 2020 · 0 repositories · arXiv:2002.03329
-
Momentum Improves Normalized SGD 9 Feb 2020 · 0 repositories · arXiv:2002.03305
-
On the distance between two neural networks and the stability of learning 9 Feb 2020 · 2 repositories · arXiv:2002.03432
-
How Good is the Bayes Posterior in Deep Neural Networks Really? 6 Feb 2020 · 1 repository · arXiv:2002.02405
-
Function approximation by neural nets in the mean-field regime: Entropic regularization and controlled McKean-Vlasov dynamics 5 Feb 2020 · 0 repositories · arXiv:2002.01987
-
Goal-Oriented Multi-Task BERT-Based Dialogue State Tracker 5 Feb 2020 · 0 repositories · arXiv:2002.02450
-
Efficient Riemannian Optimization on the Stiefel Manifold via the Cayley Transform 4 Feb 2020 · 0 repositories · arXiv:2002.01113
-
Improving Efficiency in Large-Scale Decentralized Distributed Training 4 Feb 2020 · 0 repositories · arXiv:2002.01119
-
On the Convergence of Stochastic Gradient Descent with Low-Rank Projections for Convex Low-Rank Matrix Problems 31 Jan 2020 · 0 repositories · arXiv:2001.11668
-
How Does BN Increase Collapsed Neural Network Filters? 30 Jan 2020 · 0 repositories · arXiv:2001.11216
-
Variance Reduction with Sparse Gradients 27 Jan 2020 · 0 repositories · arXiv:2001.09623
-
Stochastic Optimization of Plain Convolutional Neural Networks with Simple methods 24 Jan 2020 · 1 repository · arXiv:2001.08856
-
Intermittent Pulling with Local Compensation for Communication-Efficient Federated Learning 22 Jan 2020 · 0 repositories · arXiv:2001.08277
-
On Last-Layer Algorithms for Classification: Decoupling Representation from Uncertainty Estimation 22 Jan 2020 · 1 repository · arXiv:2001.08049Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Stochastic Item Descent Method for Large Scale Equal Circle Packing Problem 22 Jan 2020 · 0 repositories · arXiv:2001.08540
-
Harmonic Convolutional Networks based on Discrete Cosine Transform 18 Jan 2020 · 1 repository · arXiv:2001.06570Syntology official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 10 harvested samples)
-
Elastic Consistency: A General Consistency Model for Distributed Stochastic Gradient Descent 16 Jan 2020 · 0 repositories · arXiv:2001.05918
-
Backward Feature Correction: How Deep Learning Performs Deep (Hierarchical) Learning 13 Jan 2020 · 0 repositories · arXiv:2001.04413
-
Choosing the Sample with Lowest Loss makes SGD Robust 10 Jan 2020 · 1 repository · arXiv:2001.03316
-
Distributionally Robust Deep Learning using Hardness Weighted Sampling 8 Jan 2020 · 1 repository · arXiv:2001.02658
-
Poly-time universality and limitations of deep learning 7 Jan 2020 · 0 repositories · arXiv:2001.02992
-
How neural networks find generalizable solutions: Self-tuned annealing in deep learning 6 Jan 2020 · 0 repositories · arXiv:2001.01678
-
Towards understanding the true loss surface of deep neural networks using random matrix theory and iterative spectral methods 1 Jan 2020 · 0 repositories
-
Training Deep Networks with Stochastic Gradient Normalized by Layerwise Adaptive Second Moments 1 Jan 2020 · 0 repositories
-
A Dynamic Sampling Adaptive-SGD Method for Machine Learning 31 Dec 2019 · 0 repositories · arXiv:1912.13357
-
Variance Reduced Local SGD with Lower Communication Complexity 30 Dec 2019 · 1 repository · arXiv:1912.12844
-
Federated Variance-Reduced Stochastic Gradient Descent with Robustness to Byzantine Attacks 29 Dec 2019 · 0 repositories · arXiv:1912.12716
-
CProp: Adaptive Learning Rate Scaling from Past Gradient Conformity 24 Dec 2019 · 1 repository · arXiv:1912.11493
-
Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural Networks 20 Dec 2019 · 0 repositories · arXiv:1912.10095
-
Second-order Information in First-order Optimization Methods 20 Dec 2019 · 0 repositories · arXiv:1912.09926
-
Optimization for deep learning: theory and algorithms 19 Dec 2019 · 0 repositories · arXiv:1912.08957
-
Gradient-based training of Gaussian Mixture Models for High-Dimensional Streaming Data 18 Dec 2019 · 1 repository · arXiv:1912.09379
-
Generative Teaching Networks: Accelerating Neural Architecture Search by Learning to Generate Synthetic Training Data 17 Dec 2019 · 3 repositories · arXiv:1912.07768
-
Parallel Restarted SPIDER -- Communication Efficient Distributed Nonconvex Optimization with Optimal Computation Complexity 12 Dec 2019 · 0 repositories · arXiv:1912.06036
-
Linear Mode Connectivity and the Lottery Ticket Hypothesis 11 Dec 2019 · 2 repositories · arXiv:1912.05671
-
SiamMan: Siamese Motion-aware Network for Visual Tracking 11 Dec 2019 · 0 repositories · arXiv:1912.05515
-
Why are Adaptive Methods Good for Attention Models? 6 Dec 2019 · 0 repositories · arXiv:1912.03194
-
An Empirical Study on the Intrinsic Privacy of SGD 5 Dec 2019 · 1 repository · arXiv:1912.02919Syntology official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 0 honoured, 0 violated, 11 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 16 harvested samples)
-
Domain-independent Dominance of Adaptive Methods 4 Dec 2019 · 1 repository · arXiv:1912.01823
-
Stochastic Variational Inference via Upper Bound 2 Dec 2019 · 0 repositories · arXiv:1912.00650