Methods › General › Stochastic Optimization › SGD › Papers, page 19
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 19 of 21: papers 1,801 to 1,900 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Predictive Local Smoothness for Stochastic Gradient Methods 23 May 2018 · 0 repositories · arXiv:1805.09386
-
Gradient Energy Matching for Distributed Asynchronous Gradient Descent 22 May 2018 · 2 repositories · arXiv:1805.08469
-
Small steps and giant leaps: Minimal Newton solvers for Deep Learning 21 May 2018 · 6 repositories · arXiv:1805.08095Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
SmoothOut: Smoothing Out Sharp Minima to Improve Generalization in Deep Learning 21 May 2018 · 1 repository · arXiv:1805.07898Syntology official: harvested, nothing ran · 0 ran · 1 unverified (of 1 harvested sample)
-
Interpolatron: Interpolation or Extrapolation Schemes to Accelerate Optimization for Deep Neural Networks 17 May 2018 · 0 repositories · arXiv:1805.06753
-
Differential Equations for Modeling Asynchronous Algorithms 8 May 2018 · 0 repositories · arXiv:1805.02991
-
Implementation of Stochastic Quasi-Newton's Method in PyTorch 7 May 2018 · 1 repository · arXiv:1805.02338
-
A Scalable Discrete-Time Survival Model for Neural Networks 2 May 2018 · 2 repositories · arXiv:1805.00917Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
k-SVRG: Variance Reduction for Large Scale Optimization 2 May 2018 · 0 repositories · arXiv:1805.00982
-
Trainability and Accuracy of Neural Networks: An Interacting Particle System Approach 2 May 2018 · 0 repositories · arXiv:1805.00915
-
Multi-representation Ensembles and Delayed SGD Updates Improve Syntax-based NMT 1 May 2018 · 0 repositories · arXiv:1805.00456
-
A Mean Field View of the Landscape of Two-Layers Neural Networks 18 Apr 2018 · 0 repositories · arXiv:1804.06561
-
Active Mini-Batch Sampling using Repulsive Point Processes 8 Apr 2018 · 1 repository · arXiv:1804.02772
-
Byzantine Stochastic Gradient Descent 23 Mar 2018 · 0 repositories · arXiv:1803.08917
-
The Convergence of Stochastic Gradient Descent in Asynchronous Shared Memory 23 Mar 2018 · 0 repositories · arXiv:1803.08841
-
Lower error bounds for the stochastic gradient descent optimization algorithm: Sharp convergence rates for slowly and fast decaying learning rates 22 Mar 2018 · 0 repositories · arXiv:1803.08600
-
Escaping Saddles with Stochastic Gradients 15 Mar 2018 · 0 repositories · arXiv:1803.05999
-
GossipGraD: Scalable Deep Learning using Gossip Communication based Asynchronous Gradient Descent 15 Mar 2018 · 0 repositories · arXiv:1803.05880
-
On the insufficiency of existing momentum schemes for Stochastic Optimization 15 Mar 2018 · 2 repositories · arXiv:1803.05591
-
Averaging Weights Leads to Wider Optima and Better Generalization 14 Mar 2018 · 17 repositories · arXiv:1803.05407Syntology official (archive's flag): 2 ran · 7 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (of 9 harvested samples) · 4 pointer-only (licence)
-
Self-Similar Epochs: Value in Arrangement 14 Mar 2018 · 0 repositories · arXiv:1803.05389
-
Model-Agnostic Private Learning via Stability 14 Mar 2018 · 0 repositories · arXiv:1803.05101
-
High-Accuracy Low-Precision Training 9 Mar 2018 · 1 repository · arXiv:1803.03383
-
Fast Convergence for Stochastic and Distributed Gradient Descent in the Interpolation Limit 8 Mar 2018 · 0 repositories · arXiv:1803.02922
-
Slow and Stale Gradients Can Win the Race: Error-Runtime Trade-offs in Distributed SGD 3 Mar 2018 · 0 repositories · arXiv:1803.01113
-
Not All Samples Are Created Equal: Deep Learning with Importance Sampling 2 Mar 2018 · 2 repositories · arXiv:1803.00942
-
The Anisotropic Noise in Stochastic Gradient Descent: Its Behavior of Escaping from Sharp Minima and Regularization Effects 1 Mar 2018 · 1 repository · arXiv:1803.00195
-
Train Feedfoward Neural Network with Layer-wise Adaptive Rate via Approximating Back-matching Propagation 27 Feb 2018 · 0 repositories · arXiv:1802.09750
-
Shampoo: Preconditioned Stochastic Tensor Optimization 26 Feb 2018 · 3 repositories · arXiv:1802.09568Syntology 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
A Walk with SGD 24 Feb 2018 · 0 repositories · arXiv:1802.08770
-
Stochastic Gradient Descent on Highly-Parallel Architectures 24 Feb 2018 · 2 repositories · arXiv:1802.08800
-
Asynchronous Byzantine Machine Learning (the case of SGD) 22 Feb 2018 · 1 repository · arXiv:1802.07928
-
The Hidden Vulnerability of Distributed Learning in Byzantium 22 Feb 2018 · 1 repository · arXiv:1802.07927
-
Generalization Error Bounds with Probabilistic Guarantee for SGD in Nonconvex Optimization 19 Feb 2018 · 0 repositories · arXiv:1802.06903
-
An Alternative View: When Does SGD Escape Local Minima? 17 Feb 2018 · 0 repositories · arXiv:1802.06175
-
The Role of Information Complexity and Randomization in Representation Learning 14 Feb 2018 · 0 repositories · arXiv:1802.05355
-
signSGD: Compressed Optimisation for Non-Convex Problems 13 Feb 2018 · 6 repositories · arXiv:1802.04434Syntology community repositories only · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
HiGrad: Uncertainty Quantification for Online Learning and Stochastic Approximation 13 Feb 2018 · 0 repositories · arXiv:1802.04876
-
SGD and Hogwild! Convergence Without the Bounded Gradients Assumption 11 Feb 2018 · 0 repositories · arXiv:1802.03801
-
A predictor-corrector method for the training of deep neural networks 19 Jan 2018 · 0 repositories · arXiv:1803.05779
-
When Does Stochastic Gradient Algorithm Work Well? 18 Jan 2018 · 0 repositories · arXiv:1801.06159
-
MXNET-MPI: Embedding MPI parallelism in Parameter Server Task Model for scaling Deep Learning 11 Jan 2018 · 0 repositories · arXiv:1801.03855
-
How To Make the Gradients Small Stochastically: Even Faster Convex and Nonconvex SGD 8 Jan 2018 · 0 repositories · arXiv:1801.02982
-
Theory of Deep Learning IIb: Optimization Properties of SGD 7 Jan 2018 · 0 repositories · arXiv:1801.02254
-
A comparison of second-order methods for deep convolutional neural networks 1 Jan 2018 · 0 repositories
-
A Painless Attention Mechanism for Convolutional Neural Networks 1 Jan 2018 · 0 repositories
-
Better Generalization by Efficient Trust Region Method 1 Jan 2018 · 0 repositories
-
Convergence rate of sign stochastic gradient descent for non-convex functions 1 Jan 2018 · 0 repositories
-
Demystifying overcomplete nonlinear auto-encoders: fast SGD convergence towards sparse representation from random initialization 1 Jan 2018 · 0 repositories
-
Faster Distributed Synchronous SGD with Weak Synchronization 1 Jan 2018 · 0 repositories
-
Fixing Weight Decay Regularization in Adam 1 Jan 2018 · 0 repositories
-
Kronecker-factored Curvature Approximations for Recurrent Neural Networks 1 Jan 2018 · 0 repositories
-
LSH-SAMPLING BREAKS THE COMPUTATIONAL CHICKEN-AND-EGG LOOP IN ADAPTIVE STOCHASTIC GRADIENT ESTIMATION 1 Jan 2018 · 0 repositories
-
Sparse Regularized Deep Neural Networks For Efficient Embedded Learning 1 Jan 2018 · 0 repositories
-
True Asymptotic Natural Gradient Optimization 22 Dec 2017 · 0 repositories · arXiv:1712.08449
-
Improving Generalization Performance by Switching from Adam to SGD 20 Dec 2017 · 6 repositories · arXiv:1712.07628
-
On the Relationship Between the OpenAI Evolution Strategy and Stochastic Gradient Descent 18 Dec 2017 · 0 repositories · arXiv:1712.06564
-
The Power of Interpolation: Understanding the Effectiveness of SGD in Modern Over-parametrized Learning 18 Dec 2017 · 0 repositories · arXiv:1712.06559
-
Deep Gradient Compression: Reducing the Communication Bandwidth for Distributed Training 5 Dec 2017 · 6 repositories · arXiv:1712.01887
-
Machine Learning with Adversaries: Byzantine Tolerant Gradient Descent 1 Dec 2017 · 2 repositories
-
Nonlinear Acceleration of Stochastic Algorithms 1 Dec 2017 · 0 repositories
-
Online to Offline Conversions, Universality and Adaptive Minibatch Sizes 1 Dec 2017 · 0 repositories
-
Stochastic Optimization with Variance Reduction for Infinite Datasets with Finite Sum Structure 1 Dec 2017 · 0 repositories
-
Neon2: Finding Local Minima via First-Order Oracles 17 Nov 2017 · 0 repositories · arXiv:1711.06673
-
Decoupled Weight Decay Regularization 14 Nov 2017 · 23 repositories · arXiv:1711.05101Syntology community repositories only · 18 ran (of which 6 constructed an object rather than computing a result; 18 with no instrument failure: 3 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 23 harvested samples) · 5 pointer-only (licence)
-
Three Factors Influencing Minima in SGD 13 Nov 2017 · 0 repositories · arXiv:1711.04623
-
Analysis of Biased Stochastic Gradient Descent Using Sequential Semidefinite Programs 3 Nov 2017 · 0 repositories · arXiv:1711.00987
-
Don't Decay the Learning Rate, Increase the Batch Size 1 Nov 2017 · 3 repositories · arXiv:1711.00489
-
Linearly convergent stochastic heavy ball method for minimizing generalization error 30 Oct 2017 · 0 repositories · arXiv:1710.10737
-
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks 30 Oct 2017 · 0 repositories · arXiv:1710.11029
-
SGD Learns Over-parameterized Networks that Provably Generalize on Linearly Separable Data 27 Oct 2017 · 0 repositories · arXiv:1710.10174
-
Improving Negative Sampling for Word Representation using Self-embedded Features 26 Oct 2017 · 0 repositories · arXiv:1710.09805
-
A Markov Chain Theory Approach to Characterizing the Minimax Optimality of Stochastic Gradient Descent (for Least Squares) 25 Oct 2017 · 0 repositories · arXiv:1710.09430
-
Stability and Generalization of Learning Algorithms that Converge to Global Optima 23 Oct 2017 · 0 repositories · arXiv:1710.08402
-
A Novel Stochastic Stratified Average Gradient Method: Convergence Rate and Its Complexity 21 Oct 2017 · 0 repositories · arXiv:1710.07783
-
Asynchronous Decentralized Parallel Stochastic Gradient Descent 18 Oct 2017 · 3 repositories · arXiv:1710.06952
-
Synkhronos: a Multi-GPU Theano Extension for Data Parallelism 11 Oct 2017 · 1 repository · arXiv:1710.04162
-
Convergence Analysis of Distributed Stochastic Gradient Descent with Shuffling 29 Sep 2017 · 0 repositories · arXiv:1709.10432
-
Probabilistic Synchronous Parallel 22 Sep 2017 · 0 repositories · arXiv:1709.07772
-
Neural Optimizer Search with Reinforcement Learning 21 Sep 2017 · 2 repositories · arXiv:1709.07417
-
A PAC-Bayesian Analysis of Randomized Learning with Application to Stochastic Gradient Descent 19 Sep 2017 · 1 repository · arXiv:1709.06617
-
Scalable Support Vector Clustering Using Budget 19 Sep 2017 · 0 repositories · arXiv:1709.06444
-
The Impact of Local Geometry and Batch Size on Stochastic Gradient Descent for Nonconvex Problems 14 Sep 2017 · 0 repositories · arXiv:1709.04718
-
Normalized Direction-preserving Adam 13 Sep 2017 · 1 repository · arXiv:1709.04546
-
Stochastic Gradient Descent: Going As Fast As Possible But Not Faster 5 Sep 2017 · 0 repositories · arXiv:1709.01427
-
Natasha 2: Faster Non-Convex Optimization Than SGD 29 Aug 2017 · 0 repositories · arXiv:1708.08694
-
Second-Order Optimization for Non-Convex Machine Learning: An Empirical Study 25 Aug 2017 · 0 repositories · arXiv:1708.07827
-
Weighted parallel SGD for distributed unbalanced-workload training system 16 Aug 2017 · 0 repositories · arXiv:1708.04801
-
Noisy Softmax: Improving the Generalization Ability of DCNN via Postponing the Early Softmax Saturation 12 Aug 2017 · 0 repositories · arXiv:1708.03769
-
Stochastic Optimization with Bandit Sampling 8 Aug 2017 · 0 repositories · arXiv:1708.02544
-
Neural Optimizer Search using Reinforcement Learning 1 Aug 2017 · 0 repositories
-
Stochastic Adaptive Quasi-Newton Methods for Minimizing Expected Values 1 Aug 2017 · 0 repositories
-
Bridging the Gap between Constant Step Size Stochastic Gradient Descent and Markov Chains 20 Jul 2017 · 0 repositories · arXiv:1707.06386
-
Block-Normalized Gradient Method: An Empirical Study for Training Deep Neural Network 16 Jul 2017 · 2 repositories · arXiv:1707.04822
-
Dual Path Networks 6 Jul 2017 · 18 repositories · arXiv:1707.01629Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Parle: parallelizing stochastic gradient descent 3 Jul 2017 · 0 repositories · arXiv:1707.00424
-
On Scalable Inference with Stochastic Gradient Descent 1 Jul 2017 · 0 repositories · arXiv:1707.00192
-
Spectrally-normalized margin bounds for neural networks 26 Jun 2017 · 1 repository · arXiv:1706.08498
-
Collaborative Deep Learning in Fixed Topology Networks 23 Jun 2017 · 0 repositories · arXiv:1706.07880
-
Gradient Diversity: a Key Ingredient for Scalable Distributed Learning 18 Jun 2017 · 0 repositories · arXiv:1706.05699