Methods › General › Stochastic Optimization › SGD › Papers, page 18
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 18 of 21: papers 1,701 to 1,800 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
FPDeep: Scalable Acceleration of CNN Training on Deeply-Pipelined FPGA Clusters 4 Jan 2019 · 0 repositories · arXiv:1901.01007
-
SGD Converges to Global Minimum in Deep Learning via Star-convex Path 2 Jan 2019 · 0 repositories · arXiv:1901.00451
-
A continuous-time analysis of distributed stochastic gradient 28 Dec 2018 · 0 repositories · arXiv:1812.10995
-
Overparameterized Nonlinear Learning: Gradient Descent Takes the Shortest Path? 25 Dec 2018 · 0 repositories · arXiv:1812.10004
-
Stochastic Doubly Robust Gradient 21 Dec 2018 · 0 repositories · arXiv:1812.08997
-
Deep Online Learning via Meta-Learning: Continual Adaptation for Model-Based RL 18 Dec 2018 · 0 repositories · arXiv:1812.07671
-
Provable limitations of deep learning 16 Dec 2018 · 0 repositories · arXiv:1812.06369
-
Stagewise Training Accelerates Convergence of Testing Error Over SGD 10 Dec 2018 · 0 repositories · arXiv:1812.03934
-
What is the Effect of Importance Weighting in Deep Learning? 8 Dec 2018 · 1 repository · arXiv:1812.03372
-
Elastic Gossip: Distributing Neural Network Training Using Gossip-like Protocols 6 Dec 2018 · 1 repository · arXiv:1812.02407
-
Towards Theoretical Understanding of Large Batch Training in Stochastic Gradient Descent 3 Dec 2018 · 0 repositories · arXiv:1812.00542
-
A Linear Speedup Analysis of Distributed Deep Learning with Sparse and Quantized Communication 1 Dec 2018 · 0 repositories
-
Exact natural gradient in deep linear networks and its application to the nonlinear case 1 Dec 2018 · 0 repositories
-
How SGD Selects the Global Minima in Over-parameterized Learning: A Dynamical Stability Perspective 1 Dec 2018 · 1 repository
-
On the Local Hessian in Back-propagation 1 Dec 2018 · 0 repositories
-
Stochastic Composite Mirror Descent: Optimal Bounds with High Probabilities 1 Dec 2018 · 0 repositories
-
Stochastic Primal-Dual Method for Empirical Risk Minimization with O(1) Per-Iteration Complexity 1 Dec 2018 · 0 repositories
-
Training Deep Models Faster with Robust, Approximate Importance Sampling 1 Dec 2018 · 0 repositories
-
Variance-Reduced Stochastic Gradient Descent on Streaming Data 1 Dec 2018 · 0 repositories
-
ESPNetv2: A Light-weight, Power Efficient, and General Purpose Convolutional Neural Network 28 Nov 2018 · 10 repositories · arXiv:1811.11431Syntology community repositories only · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Stochastic Gradient Push for Distributed Deep Learning 27 Nov 2018 · 3 repositories · arXiv:1811.10792Syntology 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 5 unverified (of 15 harvested samples) · 2 pointer-only (licence)
-
The promises and pitfalls of Stochastic Gradient Langevin Dynamics 25 Nov 2018 · 0 repositories · arXiv:1811.10072
-
Hydra: A Peer to Peer Distributed Training & Data Collection Framework 24 Nov 2018 · 1 repository · arXiv:1811.09878
-
HyperAdam: A Learnable Task-Adaptive Adam for Network Training 22 Nov 2018 · 2 repositories · arXiv:1811.08996
-
Deep Frank-Wolfe For Neural Network Optimization 19 Nov 2018 · 1 repository · arXiv:1811.07591Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples)
-
Minimum weight norm models do not always generalize well for over-parameterized problems 16 Nov 2018 · 0 repositories · arXiv:1811.07055
-
Learning and Generalization in Overparameterized Neural Networks, Going Beyond Two Layers 12 Nov 2018 · 0 repositories · arXiv:1811.04918
-
New Convergence Aspects of Stochastic Gradient Algorithms 10 Nov 2018 · 0 repositories · arXiv:1811.12403
-
A Convergence Theory for Deep Learning via Over-Parameterization 9 Nov 2018 · 0 repositories · arXiv:1811.03962
-
On exponential convergence of SGD in non-convex over-parametrized learning 6 Nov 2018 · 0 repositories · arXiv:1811.02564
-
Deep Reinforcement Learning via L-BFGS Optimization 6 Nov 2018 · 0 repositories · arXiv:1811.02693
-
Stochastic Modified Equations and Dynamics of Stochastic Gradient Algorithms I: Mathematical Foundations 5 Nov 2018 · 0 repositories · arXiv:1811.01558
-
Stochastic Primal-Dual Method for Empirical Risk Minimization with 𝒪(1) Per-Iteration Complexity 3 Nov 2018 · 0 repositories · arXiv:1811.01182
-
Implicit Regularization of Stochastic Gradient Descent in Natural Language Processing: Observations and Implications 1 Nov 2018 · 0 repositories · arXiv:1811.00659
-
Accelerating SGD with momentum for over-parameterized learning 31 Oct 2018 · 1 repository · arXiv:1810.13395
-
Kalman Gradient Descent: Adaptive Variance Reduction in Stochastic Optimization 29 Oct 2018 · 1 repository · arXiv:1810.12273
-
On the Convergence Rate of Training Recurrent Neural Networks 29 Oct 2018 · 0 repositories · arXiv:1810.12065
-
Finding Mixed Nash Equilibria of Generative Adversarial Networks 23 Oct 2018 · 0 repositories · arXiv:1811.02002
-
ensmallen: a flexible C++ library for efficient function optimization 22 Oct 2018 · 1 repository · arXiv:1810.09361
-
Optimality of the final model found via Stochastic Gradient Descent 22 Oct 2018 · 0 repositories · arXiv:1810.09418
-
Adaptive Communication Strategies to Achieve the Best Error-Runtime Trade-off in Local-Update SGD 19 Oct 2018 · 0 repositories · arXiv:1810.08313
-
Exchangeability and Kernel Invariance in Trained MLPs 19 Oct 2018 · 0 repositories · arXiv:1810.08351
-
Small ReLU networks are powerful memorizers: a tight analysis of memorization capacity 17 Oct 2018 · 0 repositories · arXiv:1810.07770
-
Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural Networks 16 Oct 2018 · 1 repository · arXiv:1810.06773
-
Fast and Faster Convergence of SGD for Over-Parameterized Models and an Accelerated Perceptron 16 Oct 2018 · 0 repositories · arXiv:1810.07288
-
Quasi-hyperbolic momentum and Adam for deep learning 16 Oct 2018 · 2 repositories · arXiv:1810.06801Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Training Deep Neural Network in Limited Precision 12 Oct 2018 · 0 repositories · arXiv:1810.05486
-
Bayesian Deep Convolutional Networks with Many Channels are Gaussian Processes 11 Oct 2018 · 0 repositories · arXiv:1810.05148
-
signSGD with Majority Vote is Communication Efficient And Fault Tolerant 11 Oct 2018 · 4 repositories · arXiv:1810.05291Syntology 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Tight Dimension Independent Lower Bound on the Expected Convergence Rate for Diminishing Step Sizes in SGD 10 Oct 2018 · 0 repositories · arXiv:1810.04723
-
Anytime Stochastic Gradient Descent: A Time to Hear from all the Workers 6 Oct 2018 · 0 repositories · arXiv:1810.02976
-
Continuous-time Models for Stochastic Optimization Algorithms 5 Oct 2018 · 1 repository · arXiv:1810.02565
-
Large batch size training of neural networks with adversarial training and second-order information 2 Oct 2018 · 1 repository · arXiv:1810.01021
-
Directional Analysis of Stochastic Gradient Descent via von Mises-Fisher Distributions in Deep learning 29 Sep 2018 · 0 repositories · arXiv:1810.00150
-
The Convergence of Sparsified Gradient Methods 27 Sep 2018 · 0 repositories · arXiv:1809.10505
-
Preconditioner on Matrix Lie Group for SGD 26 Sep 2018 · 2 repositories · arXiv:1809.10232Syntology official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 6 pointer-only (licence)
-
Sparsified SGD with Memory 20 Sep 2018 · 1 repository · arXiv:1809.07599Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Graph-Dependent Implicit Regularisation for Distributed Stochastic Subgradient Descent 18 Sep 2018 · 0 repositories · arXiv:1809.06958
-
Discovering Low-Precision Networks Close to Full-Precision Networks for Efficient Embedded Inference 11 Sep 2018 · 0 repositories · arXiv:1809.04191
-
Learning Rate Adaptation for Federated and Differentially Private Learning 11 Sep 2018 · 1 repository · arXiv:1809.03832Syntology 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Privacy-Preserving Deep Learning via Weight Transmission 10 Sep 2018 · 0 repositories · arXiv:1809.03272
-
Stochastic Gradient Descent Learns State Equations with Nonlinear Activations 9 Sep 2018 · 0 repositories · arXiv:1809.03019
-
Decentralized Differentially Private Without-Replacement Stochastic Gradient Descent 8 Sep 2018 · 0 repositories · arXiv:1809.02727
-
Online ICA: Understanding Global Dynamics of Nonconvex Optimization via Diffusion Processes 29 Aug 2018 · 0 repositories · arXiv:1808.09642
-
Cooperative SGD: A unified Framework for the Design and Analysis of Communication-Efficient SGD Algorithms 22 Aug 2018 · 0 repositories · arXiv:1808.07576
-
Don't Use Large Mini-Batches, Use Local SGD 22 Aug 2018 · 2 repositories · arXiv:1808.07217
-
Universal Stagewise Learning for Non-Convex Problems with Convergence on Averaged Solutions 20 Aug 2018 · 0 repositories · arXiv:1808.06296
-
Ensemble Kalman Inversion: A Derivative-Free Technique For Machine Learning Tasks 10 Aug 2018 · 0 repositories · arXiv:1808.03620
-
A Unified Analysis of AdaGrad with Weighted Aggregation and Momentum Acceleration 10 Aug 2018 · 0 repositories · arXiv:1808.03408
-
Learning Overparameterized Neural Networks via Stochastic Gradient Descent on Structured Data 3 Aug 2018 · 0 repositories · arXiv:1808.01204
-
Stochastic Gradient Descent with Biased but Consistent Gradient Estimators 31 Jul 2018 · 1 repository · arXiv:1807.11880
-
A Surprising Linear Relationship Predicts Test Performance in Deep Networks 25 Jul 2018 · 3 repositories · arXiv:1807.09659
-
signProx: One-Bit Proximal Algorithm for Nonconvex Stochastic Optimization 20 Jul 2018 · 0 repositories · arXiv:1807.08023
-
Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods 19 Jul 2018 · 1 repository · arXiv:1807.07540
-
Parallel Restarted SGD with Faster Convergence and Less Communication: Demystifying Why Model Averaging Works for Deep Learning 17 Jul 2018 · 0 repositories · arXiv:1807.06629
-
Evolving Differentiable Gene Regulatory Networks 16 Jul 2018 · 0 repositories · arXiv:1807.05948
-
On the Relation Between the Sharpest Directions of DNN Loss and the SGD Step Length 13 Jul 2018 · 1 repository · arXiv:1807.05031
-
Maximizing Invariant Data Perturbation with Stochastic Optimization 12 Jul 2018 · 1 repository · arXiv:1807.05077
-
Metalearning with Hebbian Fast Weights 12 Jul 2018 · 0 repositories · arXiv:1807.05076
-
Quasi-Monte Carlo Variational Inference 4 Jul 2018 · 0 repositories · arXiv:1807.01604
-
Batch IS NOT Heavy: Learning Word Representations From All Samples 1 Jul 2018 · 0 repositories
-
Random Shuffling Beats SGD after Finite Epochs 26 Jun 2018 · 0 repositories · arXiv:1806.10077
-
Faster SGD training by minibatch persistency 19 Jun 2018 · 0 repositories · arXiv:1806.07353
-
Closing the Generalization Gap of Adaptive Gradient Methods in Training Deep Neural Networks 18 Jun 2018 · 2 repositories · arXiv:1806.06763
-
Using Mode Connectivity for Loss Landscape Analysis 18 Jun 2018 · 0 repositories · arXiv:1806.06977
-
There Are Many Consistent Explanations of Unlabeled Data: Why You Should Average 14 Jun 2018 · 2 repositories · arXiv:1806.05594
-
Boosted Training of Convolutional Neural Networks for Multi-Class Segmentation 13 Jun 2018 · 0 repositories · arXiv:1806.05974
-
When Will Gradient Methods Converge to Max-margin Classifier under ReLU Models? 12 Jun 2018 · 1 repository · arXiv:1806.04339
-
Full deep neural network training on a pruned weight budget 11 Jun 2018 · 1 repository · arXiv:1806.06949
-
The Effect of Network Width on the Performance of Large-batch Training 11 Jun 2018 · 0 repositories · arXiv:1806.03791
-
Lightweight Stochastic Optimization for Minimizing Finite Sums with Infinite Data 8 Jun 2018 · 0 repositories · arXiv:1806.02927
-
Probabilistic Deep Learning using Random Sum-Product Networks 5 Jun 2018 · 0 repositories · arXiv:1806.01910
-
Stochastic Gradient Descent on Separable Data: Exact Convergence with a Fixed Learning Rate 5 Jun 2018 · 0 repositories · arXiv:1806.01796
-
Backdrop: Stochastic Backpropagation 4 Jun 2018 · 1 repository · arXiv:1806.01337
-
Geometry Aware Constrained Optimization Techniques for Deep Learning 1 Jun 2018 · 0 repositories
-
On Consensus-Optimality Trade-offs in Collaborative Deep Learning 30 May 2018 · 0 repositories · arXiv:1805.12120
-
How Much Restricted Isometry is Needed In Nonconvex Matrix Recovery? 25 May 2018 · 0 repositories · arXiv:1805.10251
-
Statistical Optimality of Stochastic Gradient Descent on Hard Learning Problems through Multiple Passes 25 May 2018 · 0 repositories · arXiv:1805.10074
-
Zeno: Distributed Stochastic Gradient Descent with Suspicion-based Fault-tolerance 25 May 2018 · 1 repository · arXiv:1805.10032
-
Local SGD Converges Fast and Communicates Little 24 May 2018 · 2 repositories · arXiv:1805.09767