Methods › General › Stochastic Optimization › SGD › Papers, page 20
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 20 of 21: papers 1,901 to 2,000 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
YellowFin and the Art of Momentum Tuning 12 Jun 2017 · 2 repositories · arXiv:1706.03471Syntology 0 ran · 1 unverified (of 1 harvested sample)
-
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour 8 Jun 2017 · 73 repositories · arXiv:1706.02677Syntology 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
CATERPILLAR: Coarse Grain Reconfigurable Architecture for Accelerating the Training of Deep Neural Networks 1 Jun 2017 · 0 repositories · arXiv:1706.00517
-
Reinforcement Learning for Learning Rate Control 31 May 2017 · 0 repositories · arXiv:1705.11159
-
Online to Offline Conversions, Universality and Adaptive Minibatch Sizes 30 May 2017 · 0 repositories · arXiv:1705.10499
-
Convergence Analysis of Two-layer Neural Networks with ReLU Activation 28 May 2017 · 0 repositories · arXiv:1705.09886
-
The Marginal Value of Adaptive Gradient Methods in Machine Learning 23 May 2017 · 3 repositories · arXiv:1705.08292
-
Dissecting Adam: The Sign, Magnitude and Variance of Stochastic Gradients 22 May 2017 · 2 repositories · arXiv:1705.07774
-
On the diffusion approximation of nonconvex stochastic gradient descent 22 May 2017 · 0 repositories · arXiv:1705.07562
-
Parallel Stochastic Gradient Descent with Sound Combiners 22 May 2017 · 0 repositories · arXiv:1705.08030
-
Statistical inference using SGD 21 May 2017 · 0 repositories · arXiv:1705.07477
-
Determinantal Point Processes for Mini-Batch Diversification 1 May 2017 · 0 repositories · arXiv:1705.00607
-
Parseval Networks: Improving Robustness to Adversarial Examples 28 Apr 2017 · 1 repository · arXiv:1704.08847
-
Linear Convergence of Accelerated Stochastic Gradient Descent for Nonconvex Nonsmooth Optimization 26 Apr 2017 · 0 repositories · arXiv:1704.07953
-
Active Bias: Training More Accurate Neural Networks by Emphasizing High Variance Samples 24 Apr 2017 · 1 repository · arXiv:1704.07433
-
Stochastic Gradient Descent as Approximate Bayesian Inference 13 Apr 2017 · 1 repository · arXiv:1704.04289
-
Computing Nonvacuous Generalization Bounds for Deep (Stochastic) Neural Networks with Many More Parameters than Training Data 31 Mar 2017 · 3 repositories · arXiv:1703.11008Syntology 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (of 3 harvested samples) · 2 pointer-only (licence)
-
Theory II: Landscape of the Empirical Risk in Deep Learning 28 Mar 2017 · 0 repositories · arXiv:1703.09833
-
Empirical Evaluation of Parallel Training Algorithms on Acoustic Modeling 17 Mar 2017 · 0 repositories · arXiv:1703.05880
-
Separation of time scales and direct computation of weights in deep neural networks 14 Mar 2017 · 0 repositories · arXiv:1703.04757
-
Data-Dependent Stability of Stochastic Gradient Descent 5 Mar 2017 · 0 repositories · arXiv:1703.01678
-
A Robust Adaptive Stochastic Gradient Method for Deep Learning 2 Mar 2017 · 1 repository · arXiv:1703.00788
-
SARAH: A Novel Method for Machine Learning Problems Using Stochastic Recursive Gradient 1 Mar 2017 · 0 repositories · arXiv:1703.00102
-
Learning What Data to Learn 28 Feb 2017 · 0 repositories · arXiv:1702.08635
-
McKernel: A Library for Approximate Kernel Expansions in Log-linear Time 27 Feb 2017 · 1 repository · arXiv:1702.08159
-
SGD Learns the Conjugate Kernel Class of the Network 27 Feb 2017 · 0 repositories · arXiv:1702.08503
-
On SGD's Failure in Practice: Characterizing and Overcoming Stalling 1 Feb 2017 · 0 repositories · arXiv:1702.00317
-
Reinforced stochastic gradient descent for deep neural network learning 27 Jan 2017 · 0 repositories · arXiv:1701.07974
-
Optimization on Product Submanifolds of Convolution Kernels 22 Jan 2017 · 0 repositories · arXiv:1701.06123
-
Efficient Distributed Semi-Supervised Learning using Stochastic Regularization over Affinity Graphs 15 Dec 2016 · 0 repositories · arXiv:1612.04898
-
Tuning the Scheduling of Distributed Stochastic Gradient Descent with Bayesian Optimization 1 Dec 2016 · 0 repositories · arXiv:1612.00383
-
Spatial contrasting for deep unsupervised learning 21 Nov 2016 · 0 repositories · arXiv:1611.06996
-
How to scale distributed deep learning? 14 Nov 2016 · 0 repositories · arXiv:1611.04581
-
Entropy-SGD: Biasing Gradient Descent Into Wide Valleys 6 Nov 2016 · 2 repositories · arXiv:1611.01838
-
Quasi-Recurrent Neural Networks 5 Nov 2016 · 7 repositories · arXiv:1611.01576Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Statistical Inference for Model Parameters in Stochastic Gradient Descent 27 Oct 2016 · 0 repositories · arXiv:1610.08637
-
Optimization on Submanifolds of Convolution Kernels in CNNs 22 Oct 2016 · 0 repositories · arXiv:1610.07008
-
An Efficient Minibatch Acceptance Test for Metropolis-Hastings 19 Oct 2016 · 0 repositories · arXiv:1610.06848
-
CuMF_SGD: Fast and Scalable Matrix Factorization 19 Oct 2016 · 2 repositories · arXiv:1610.05838
-
Big Batch SGD: Automated Inference using Adaptive Batch Sizes 18 Oct 2016 · 0 repositories · arXiv:1610.05792
-
Parallelizing Stochastic Gradient Descent for Least Squares Regression: mini-batching, averaging, and model misspecification 12 Oct 2016 · 4 repositories · arXiv:1610.03774
-
QSGD: Communication-Efficient SGD via Gradient Quantization and Encoding 7 Oct 2016 · 2 repositories · arXiv:1610.02132Syntology 0 ran · 3 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
Near-Data Processing for Differentiable Machine Learning Models 6 Oct 2016 · 0 repositories · arXiv:1610.02273
-
Stochastic Optimization with Variance Reduction for Infinite Datasets with Finite-Sum Structure 4 Oct 2016 · 1 repository · arXiv:1610.00970
-
Deep unsupervised learning through spatial contrasting 2 Oct 2016 · 0 repositories · arXiv:1610.00243
-
Asynchronous Stochastic Gradient Descent with Delay Compensation 27 Sep 2016 · 0 repositories · arXiv:1609.08326
-
Generalization Error Bounds for Optimization Algorithms via Stability 27 Sep 2016 · 0 repositories · arXiv:1609.08397
-
Data Dependent Convergence for Distributed Stochastic Optimization 30 Aug 2016 · 0 repositories · arXiv:1608.08337
-
Lets keep it simple, Using simple architectures to outperform deeper and more complex architectures 22 Aug 2016 · 9 repositories · arXiv:1608.06037Syntology community repositories only · 11 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 11 harvested samples)
-
Uniform Generalization, Concentration, and Adaptive Learning 22 Aug 2016 · 0 repositories · arXiv:1608.06072
-
Parallel SGD: When does averaging help? 23 Jun 2016 · 0 repositories · arXiv:1606.07365
-
Bolt-on Differential Privacy for Scalable Stochastic Gradient Descent-based Analytics 15 Jun 2016 · 1 repository · arXiv:1606.04722
-
Distributed Hessian-Free Optimization for Deep Neural Network 2 Jun 2016 · 0 repositories · arXiv:1606.00511
-
Asynchrony begets Momentum, with an Application to Deep Learning 31 May 2016 · 3 repositories · arXiv:1605.09774
-
CYCLADES: Conflict-free Asynchronous Machine Learning 31 May 2016 · 1 repository · arXiv:1605.09721
-
Provable Efficient Online Matrix Completion via Non-convex Stochastic Gradient Descent 26 May 2016 · 0 repositories · arXiv:1605.08370
-
Path-Normalized Optimization of Recurrent Neural Networks with ReLU Activations 23 May 2016 · 0 repositories · arXiv:1605.07154
-
Barzilai-Borwein Step Size for Stochastic Gradient Descent 13 May 2016 · 1 repository · arXiv:1605.04131Syntology 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 7 harvested samples)
-
Distributed stochastic optimization for deep learning (thesis) 7 May 2016 · 0 repositories · arXiv:1605.02216
-
On the Convergence of A Family of Robust Losses for Stochastic Gradient Descent 5 May 2016 · 0 repositories · arXiv:1605.01623
-
Greedy Criterion in Orthogonal Greedy Learning 20 Apr 2016 · 0 repositories · arXiv:1604.05993
-
Asynchronous Stochastic Gradient Descent with Variance Reduction for Non-Convex Optimization 12 Apr 2016 · 0 repositories · arXiv:1604.03584
-
Efficient Globally Convergent Stochastic Optimization for Canonical Correlation Analysis 7 Apr 2016 · 0 repositories · arXiv:1604.01870
-
Stochastic Variance Reduction for Nonconvex Optimization 19 Mar 2016 · 0 repositories · arXiv:1603.06160
-
Accelerating Deep Neural Network Training with Inconsistent Stochastic Gradient Descent 17 Mar 2016 · 0 repositories · arXiv:1603.05544
-
Distributed Deep Learning Using Synchronous Stochastic Gradient Descent 22 Feb 2016 · 0 repositories · arXiv:1602.06709
-
A Variational Analysis of Stochastic Gradient Algorithms 8 Feb 2016 · 0 repositories · arXiv:1602.02666
-
Loss factorization, weakly supervised learning and label noise robustness 8 Feb 2016 · 0 repositories · arXiv:1602.02450
-
A Kronecker-factored approximate Fisher matrix for convolution layers 3 Feb 2016 · 4 repositories · arXiv:1602.01407Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Training Recurrent Neural Networks by Diffusion 16 Jan 2016 · 0 repositories · arXiv:1601.04114
-
A Light Touch for Heavily Constrained SGD 15 Dec 2015 · 0 repositories · arXiv:1512.04960
-
Preconditioned Stochastic Gradient Descent 14 Dec 2015 · 2 repositories · arXiv:1512.04202Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Efficient Distributed SGD with Variance Reduction 9 Dec 2015 · 0 repositories · arXiv:1512.02970
-
Variance Reduction for Distributed Stochastic Gradient Descent 5 Dec 2015 · 0 repositories · arXiv:1512.01708
-
Kalman-based Stochastic Gradient Method with Stop Condition and Insensitivity to Conditioning 3 Dec 2015 · 0 repositories · arXiv:1512.01139
-
Finite-Time Analysis of Projected Langevin Monte Carlo 1 Dec 2015 · 0 repositories
-
Taming the Wild: A Unified Analysis of Hogwild-Style Algorithms 1 Dec 2015 · 0 repositories
-
Distributed Deep Learning for Question Answering 3 Nov 2015 · 0 repositories · arXiv:1511.01158
-
How Important is Weight Symmetry in Backpropagation? 17 Oct 2015 · 2 repositories · arXiv:1510.05067
-
Uniform Learning in a Deep Neural Network via "Oddball" Stochastic Gradient Descent 8 Oct 2015 · 0 repositories · arXiv:1510.02442
-
Convergence of Stochastic Gradient Descent for PCA 30 Sep 2015 · 0 repositories · arXiv:1509.09002
-
Stochastic gradient descent methods for estimation with large data sets 22 Sep 2015 · 1 repository · arXiv:1509.06459
-
"Oddball SGD": Novelty Driven Stochastic Gradient Descent for Training Deep Neural Networks 18 Sep 2015 · 0 repositories · arXiv:1509.05765
-
Fast Asynchronous Parallel Stochastic Gradient Decent 24 Aug 2015 · 0 repositories · arXiv:1508.05711
-
On the Convergence of SGD Training of Neural Networks 12 Aug 2015 · 0 repositories · arXiv:1508.02790
-
Beyond Convexity: Stochastic Quasi-Convex Optimization 8 Jul 2015 · 0 repositories · arXiv:1507.02030
-
Online Learning to Sample 30 Jun 2015 · 0 repositories · arXiv:1506.09016
-
Stochastic Gradient Made Stable: A Manifold Propagation Approach for Large-Scale Optimization 28 Jun 2015 · 0 repositories · arXiv:1506.08350
-
On Variance Reduction in Stochastic Gradient Descent and its Asynchronous Variants 23 Jun 2015 · 0 repositories · arXiv:1506.06840
-
Taming the Wild: A Unified Analysis of Hogwild!-Style Algorithms 22 Jun 2015 · 0 repositories · arXiv:1506.06438
-
Variance Reduced Stochastic Gradient Descent with Neighbors 11 Jun 2015 · 0 repositories · arXiv:1506.03662
-
Path-SGD: Path-Normalized Optimization in Deep Neural Networks 8 Jun 2015 · 1 repository · arXiv:1506.02617
-
Spatial Transformer Networks 5 Jun 2015 · 45 repositories · arXiv:1506.02025Syntology 21 ran (of which 0 constructed an object rather than computing a result; 21 with no instrument failure: 1 honoured, 0 violated, 20 with no contract checked; 0 where Syntology's instrument failed) · 10 unverified (of 31 harvested samples)
-
Early Stopping is Nonparametric Variational Inference 6 Apr 2015 · 1 repository · arXiv:1504.01344
-
Learning Co-Sparse Analysis Operators with Separable Structures 9 Mar 2015 · 0 repositories · arXiv:1503.02398
-
ADASECANT: Robust Adaptive Secant Method for Stochastic Gradient 23 Dec 2014 · 0 repositories · arXiv:1412.7419
-
Learning from Data with Heterogeneous Noise using SGD 17 Dec 2014 · 0 repositories · arXiv:1412.5617
-
The Loss Surfaces of Multilayer Networks 30 Nov 2014 · 1 repository · arXiv:1412.0233
-
Global Convergence of Stochastic Gradient Descent for Some Non-convex Matrix Problems 5 Nov 2014 · 0 repositories · arXiv:1411.1134
-
Parallel training of DNNs with Natural Gradient and Parameter Averaging 27 Oct 2014 · 1 repository · arXiv:1410.7455