Methods › General › Stochastic Optimization › SGD › Papers, page 17
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 17 of 21: papers 1,601 to 1,700 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Adaptively Preconditioned Stochastic Gradient Langevin Dynamics 10 Jun 2019 · 1 repository · arXiv:1906.04324
-
Stochastic Mirror Descent on Overparameterized Nonlinear Models: Convergence, Implicit Regularization, and Generalization 10 Jun 2019 · 1 repository · arXiv:1906.03830
-
Making Asynchronous Stochastic Gradient Descent Work for Transformers 8 Jun 2019 · 0 repositories · arXiv:1906.03496
-
Bad Global Minima Exist and SGD Can Reach Them 6 Jun 2019 · 1 repository · arXiv:1906.02613Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Qsparse-local-SGD: Distributed SGD with Quantization, Sparsification, and Local Computations 6 Jun 2019 · 0 repositories · arXiv:1906.02367
-
Binarized Collaborative Filtering with Distilling Graph Convolutional Networks 5 Jun 2019 · 0 repositories · arXiv:1906.01829
-
Embedded hyper-parameter tuning by Simulated Annealing 4 Jun 2019 · 2 repositories · arXiv:1906.01504
-
PowerSGD: Practical Low-Rank Gradient Compression for Distributed Optimization 31 May 2019 · 1 repository · arXiv:1905.13727Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Global Momentum Compression for Sparse Communication in Distributed Learning 30 May 2019 · 0 repositories · arXiv:1905.12948
-
On the Convergence of Memory-Based Distributed SGD 30 May 2019 · 0 repositories · arXiv:1905.12960
-
P3SGD: Patient Privacy Preserving SGD for Regularizing Deep CNNs in Pathological Image Classification 30 May 2019 · 0 repositories · arXiv:1905.12883
-
Accelerated Sparsified SGD with Error Feedback 29 May 2019 · 0 repositories · arXiv:1905.12224
-
Privacy Amplification by Mixing and Diffusion Mechanisms 29 May 2019 · 0 repositories · arXiv:1905.12264
-
Gram-Gauss-Newton Method: Learning Overparameterized Neural Networks for Regression Problems 28 May 2019 · 0 repositories · arXiv:1905.11675
-
SGD on Neural Networks Learns Functions of Increasing Complexity 28 May 2019 · 1 repository · arXiv:1905.11604
-
Communication-Efficient Distributed Blockwise Momentum SGD with Error-Feedback 27 May 2019 · 1 repository · arXiv:1905.10936
-
Natural Compression for Distributed Deep Learning 27 May 2019 · 0 repositories · arXiv:1905.10988
-
Stochastic Gradient Methods with Layer-wise Adaptive Moments for Training of Deep Networks 27 May 2019 · 3 repositories · arXiv:1905.11286Syntology 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 3 unverified (of 16 harvested samples) · 1 pointer-only (licence)
-
Stochastic Shared Embeddings: Data-driven Regularization of Embedding Layers 25 May 2019 · 3 repositories · arXiv:1905.10630
-
VecHGrad for Solving Accurately Complex Tensor Decomposition 24 May 2019 · 0 repositories · arXiv:1905.12413
-
Painless Stochastic Gradient: Interpolation, Line-Search, and Convergence Rates 24 May 2019 · 1 repository · arXiv:1905.09997
-
MATCHA: Speeding Up Decentralized SGD via Matching Decomposition Sampling 23 May 2019 · 4 repositories · arXiv:1905.09435Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples)
-
Fine-grained Optimization of Deep Neural Networks 22 May 2019 · 1 repository · arXiv:1905.09054
-
Time-Smoothed Gradients for Online Forecasting 21 May 2019 · 0 repositories · arXiv:1905.08850
-
Shaping the learning landscape in neural networks around wide flat minima 20 May 2019 · 0 repositories · arXiv:1905.07833
-
Adaptively Truncating Backpropagation Through Time to Control Gradient Bias 17 May 2019 · 1 repository · arXiv:1905.07473Syntology official (archive's flag): 18 ran · 18 ran (of which 0 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 0 violated, 18 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 21 harvested samples)
-
Meta Reinforcement Learning with Task Embedding and Shared Policy 16 May 2019 · 2 repositories · arXiv:1905.06527
-
DoubleSqueeze: Parallel Stochastic Gradient Descent with Double-Pass Error-Compensated Compression 15 May 2019 · 0 repositories · arXiv:1905.05957
-
On the Computation and Communication Complexity of Parallel SGD with Dynamic Batch Sizes for Stochastic Non-Convex Optimization 10 May 2019 · 0 repositories · arXiv:1905.04346
-
The Effect of Network Width on Stochastic Gradient Descent and Generalization: an Empirical Study 9 May 2019 · 0 repositories · arXiv:1905.03776
-
On the Linear Speedup Analysis of Communication Efficient Momentum SGD for Distributed Non-Convex Optimization 9 May 2019 · 0 repositories · arXiv:1905.03817
-
AutoAssist: A Framework to Accelerate Training of Deep Neural Networks 8 May 2019 · 1 repository · arXiv:1905.03381Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Fast and Robust Distributed Learning in High Dimension 5 May 2019 · 0 repositories · arXiv:1905.04374
-
Processing Megapixel Images with Deep Attention-Sampling Models 3 May 2019 · 2 repositories · arXiv:1905.03711Syntology community repositories only · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 6 harvested samples)
-
A unified theory of adaptive stochastic gradient descent as Bayesian filtering 1 May 2019 · 0 repositories
-
A Walk with SGD: How SGD Explores Regions of Deep Network Loss? 1 May 2019 · 0 repositories
-
Asynchronous SGD without gradient delay for efficient distributed training 1 May 2019 · 0 repositories
-
DANA: Scalable Out-of-the-box Distributed ASGD Without Retuning 1 May 2019 · 0 repositories
-
Distributionally Robust Optimization Leads to Better Generalization: on SGD and Beyond 1 May 2019 · 0 repositories
-
G-SGD: Optimizing ReLU Neural Networks in its Positively Scale-Invariant Space 1 May 2019 · 0 repositories
-
Learn From Neighbour: A Curriculum That Train Low Weighted Samples By Imitating 1 May 2019 · 0 repositories
-
LSH Microbatches for Stochastic Gradients: Value in Rearrangement 1 May 2019 · 0 repositories
-
On the Trajectory of Stochastic Gradient Descent in the Information Plane 1 May 2019 · 0 repositories
-
Online Hyperparameter Adaptation via Amortized Proximal Optimization 1 May 2019 · 0 repositories
-
Padam: Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks 1 May 2019 · 0 repositories
-
The Anisotropic Noise in Stochastic Gradient Descent: Its Behavior of Escaping from Minima and Regularization Effects 1 May 2019 · 0 repositories
-
Making the Last Iterate of SGD Information Theoretically Optimal 29 Apr 2019 · 0 repositories · arXiv:1904.12443
-
The Step Decay Schedule: A Near Optimal, Geometrically Decaying Learning Rate Procedure For Least Squares 29 Apr 2019 · 1 repository · arXiv:1904.12838Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources 26 Apr 2019 · 2 repositories · arXiv:1904.12043Syntology 13 ran (of which 0 constructed an object rather than computing a result; 13 with no instrument failure: 0 honoured, 0 violated, 13 with no contract checked; 0 where Syntology's instrument failed) · 7 unverified (of 20 harvested samples) · 14 pointer-only (licence)
-
SWALP : Stochastic Weight Averaging in Low-Precision Training 26 Apr 2019 · 3 repositories · arXiv:1904.11943Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
Communication trade-offs for synchronized distributed SGD with large step size 25 Apr 2019 · 0 repositories · arXiv:1904.11325
-
RepPoints: Point Set Representation for Object Detection 25 Apr 2019 · 6 repositories · arXiv:1904.11490Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Stability and Optimization Error of Stochastic Gradient Descent for Pairwise Learning 25 Apr 2019 · 0 repositories · arXiv:1904.11316
-
Semi-Cyclic Stochastic Gradient Descent 23 Apr 2019 · 0 repositories · arXiv:1904.10120
-
Distributed Deep Learning Strategies For Automatic Speech Recognition 10 Apr 2019 · 0 repositories · arXiv:1904.04956
-
Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution 10 Apr 2019 · 28 repositories · arXiv:1904.05049Syntology community repositories only · 24 ran (of which 0 constructed an object rather than computing a result; 11 with no instrument failure: 1 honoured, 0 violated, 10 with no contract checked; 13 where Syntology's instrument failed) · 10 unverified (of 34 harvested samples) · 9 pointer-only (licence)
-
Centripetal SGD for Pruning Very Deep Convolutional Networks with Complicated Structure 8 Apr 2019 · 1 repository · arXiv:1904.03837
-
A Stochastic Interpretation of Stochastic Mirror Descent: Risk-Sensitive Optimality 3 Apr 2019 · 0 repositories · arXiv:1904.01855
-
Exponentially convergent stochastic k-PCA without variance reduction 3 Apr 2019 · 0 repositories · arXiv:1904.01750
-
Normal Approximation for Stochastic Gradient Descent via Non-Asymptotic Rates of Martingale CLT 3 Apr 2019 · 0 repositories · arXiv:1904.02130
-
Lessons from Building Acoustic Models with a Million Hours of Speech 2 Apr 2019 · 0 repositories · arXiv:1904.01624
-
On the Stability and Generalization of Learning with Kernel Activation Functions 28 Mar 2019 · 0 repositories · arXiv:1903.11990
-
Learning Competitive and Discriminative Reconstructions for Anomaly Detection 17 Mar 2019 · 0 repositories · arXiv:1903.07058
-
Inefficiency of K-FAC for Large Batch Size Training 14 Mar 2019 · 0 repositories · arXiv:1903.06237
-
A Distributed Hierarchical SGD Algorithm with Sparse Global Reduction 12 Mar 2019 · 0 repositories · arXiv:1903.05133
-
Communication-efficient distributed SGD with Sketching 12 Mar 2019 · 2 repositories · arXiv:1903.04488Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Accelerating Minibatch Stochastic Gradient Descent using Typicality Sampling 11 Mar 2019 · 0 repositories · arXiv:1903.04192
-
Partially Shuffling the Training Data to Improve Language Models 11 Mar 2019 · 1 repository · arXiv:1903.04167
-
Fall of Empires: Breaking Byzantine-tolerant SGD by Inner Product Manipulation 10 Mar 2019 · 4 repositories · arXiv:1903.03936Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Time-Delay Momentum: A Regularization Perspective on the Convergence and Generalization of Stochastic Momentum for Deep Learning 2 Mar 2019 · 0 repositories · arXiv:1903.00760
-
Distributed Byzantine Tolerant Stochastic Gradient Descent in the Era of Big Data 27 Feb 2019 · 0 repositories · arXiv:1902.10336
-
Equi-normalization of Neural Networks 27 Feb 2019 · 1 repository · arXiv:1902.10416
-
Adaptive Gradient Methods with Dynamic Bound of Learning Rate 26 Feb 2019 · 5 repositories · arXiv:1902.09843
-
Beating SGD Saturation with Tail-Averaging and Minibatching 22 Feb 2019 · 0 repositories · arXiv:1902.08668
-
Optimizing Stochastic Gradient Descent in Text Classification Based on Fine-Tuning Hyper-Parameters Approach. A Case Study on Automatic Classification of Global Terrorist Attacks 18 Feb 2019 · 0 repositories · arXiv:1902.06542
-
MultiGrain: a unified image embedding for classes and instances 14 Feb 2019 · 3 repositories · arXiv:1902.05509Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Scaling Limits of Wide Neural Networks with Weight Sharing: Gaussian Process Behavior, Gradient Independence, and Neural Tangent Kernel Derivation 13 Feb 2019 · 0 repositories · arXiv:1902.04760
-
On Nonconvex Optimization for Machine Learning: Gradients, Stochasticity, and Saddle Points 13 Feb 2019 · 0 repositories · arXiv:1902.04811
-
A Simple Baseline for Bayesian Uncertainty in Deep Learning 7 Feb 2019 · 8 repositories · arXiv:1902.02476Syntology official (archive's flag): 8 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 6 unverified (of 17 harvested samples) · 9 pointer-only (licence)
-
A Scale Invariant Flatness Measure for Deep Network Minima 6 Feb 2019 · 0 repositories · arXiv:1902.02434
-
The role of a layer in deep neural networks: a Gaussian Process perspective 6 Feb 2019 · 0 repositories · arXiv:1902.02354
-
Distribution-Dependent Analysis of Gibbs-ERM Principle 5 Feb 2019 · 0 repositories · arXiv:1902.01846
-
Stochastic Gradient Descent for Nonconvex Learning without Bounded Gradient Assumptions 3 Feb 2019 · 0 repositories · arXiv:1902.00908
-
Asymmetric Valleys: Beyond Sharp and Flat Local Minima 2 Feb 2019 · 1 repository · arXiv:1902.00744Syntology 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
Uniform-in-Time Weak Error Analysis for Stochastic Gradient Descent Algorithms via Diffusion Approximation 2 Feb 2019 · 0 repositories · arXiv:1902.00635
-
Compressing Gradient Optimizers via Count-Sketches 1 Feb 2019 · 1 repository · arXiv:1902.00179
-
Sharp Analysis for Nonconvex SGD Escaping from Saddle Points 1 Feb 2019 · 0 repositories · arXiv:1902.00247
-
Error Feedback Fixes SignSGD and other Gradient Compression Schemes 28 Jan 2019 · 2 repositories · arXiv:1901.09847Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
ICLR Reproducibility Challenge Report (Padam : Closing The Generalization Gap Of Adaptive Gradient Methods in Training Deep Neural Networks) 28 Jan 2019 · 1 repository · arXiv:1901.09517
-
99% of Distributed Optimization is a Waste of Time: The Issue and How to Fix it 27 Jan 2019 · 0 repositories · arXiv:1901.09437
-
Augment your batch: better training with larger batches 27 Jan 2019 · 1 repository · arXiv:1901.09335
-
SGD: General Analysis and Improved Rates 27 Jan 2019 · 0 repositories · arXiv:1901.09401
-
Escaping Saddle Points with Adaptive Gradient Methods 26 Jan 2019 · 0 repositories · arXiv:1901.09149
-
Generalisation dynamics of online learning in over-parameterised neural networks 25 Jan 2019 · 0 repositories · arXiv:1901.09085
-
Surrogate Losses for Online Learning of Stepsizes in Stochastic Non-Convex Optimization 25 Jan 2019 · 1 repository · arXiv:1901.09068
-
Fitting ReLUs via SGD and Quantized SGD 19 Jan 2019 · 0 repositories · arXiv:1901.06587
-
A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks 18 Jan 2019 · 1 repository · arXiv:1901.06053
-
Quasi-potential as an implicit regularizer for the loss function in the stochastic gradient descent 18 Jan 2019 · 0 repositories · arXiv:1901.06054
-
Recombination of Artificial Neural Networks 12 Jan 2019 · 0 repositories · arXiv:1901.03900
-
Quantized Epoch-SGD for Communication-Efficient Distributed Learning 10 Jan 2019 · 0 repositories · arXiv:1901.03040