Methods › General › Stochastic Optimization › SGD › Papers, page 16
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 16 of 21: papers 1,501 to 1,600 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Communication trade-offs for Local-SGD with large step size 1 Dec 2019 · 1 repository
-
Control Batch Size and Learning Rate to Generalize Well: Theoretical and Empirical Evidence 1 Dec 2019 · 0 repositories
-
Efficient Meta Learning via Minibatch Proximal Update 1 Dec 2019 · 0 repositories
-
Fast and Accurate Stochastic Gradient Estimation 1 Dec 2019 · 1 repository
-
Optimal Sparsity-Sensitive Bounds for Distributed Mean Estimation 1 Dec 2019 · 1 repository
-
Qsparse-local-SGD: Distributed SGD with Quantization, Sparsification and Local Computations 1 Dec 2019 · 1 repository
-
A Hybrid Approach Towards Two Stage Bengali Question Classification Utilizing Smart Data Balancing Technique 30 Nov 2019 · 0 repositories · arXiv:1912.00127
-
On the Heavy-Tailed Theory of Stochastic Gradient Descent for Deep Neural Networks 29 Nov 2019 · 0 repositories · arXiv:1912.00018
-
Non-parametric Uni-modality Constraints for Deep Ordinal Classification 25 Nov 2019 · 1 repository · arXiv:1911.10720
-
Neural Networks Learning and Memorization with (almost) no Over-Parameterization 22 Nov 2019 · 0 repositories · arXiv:1911.09873
-
Parameter-Free Locally Differentially Private Stochastic Subgradient Descent 21 Nov 2019 · 0 repositories · arXiv:1911.09564
-
Bayesian interpretation of SGD as Ito process 20 Nov 2019 · 0 repositories · arXiv:1911.09011
-
Local AdaAlter: Communication-Efficient Stochastic Gradient Descent with Adaptive Learning Rates 20 Nov 2019 · 1 repository · arXiv:1911.09030
-
Optimal Mini-Batch Size Selection for Fast Gradient Descent 15 Nov 2019 · 0 repositories · arXiv:1911.06459
-
Throughput Prediction of Asynchronous SGD in TensorFlow 12 Nov 2019 · 0 repositories · arXiv:1911.04650
-
MindTheStep-AsyncPSGD: Adaptive Asynchronous Parallel Stochastic Gradient Descent 8 Nov 2019 · 0 repositories · arXiv:1911.03444
-
A Rule for Gradient Estimator Selection, with an Application to Variational Inference 5 Nov 2019 · 0 repositories · arXiv:1911.01894
-
Reinforced Product Metadata Selection for Helpfulness Assessment of Customer Reviews 1 Nov 2019 · 0 repositories
-
Mixing of Stochastic Accelerated Gradient Descent 31 Oct 2019 · 0 repositories · arXiv:1910.14616
-
On the Convergence of Local Descent Methods in Federated Learning 31 Oct 2019 · 0 repositories · arXiv:1910.14425
-
Local SGD with Periodic Averaging: Tighter Analysis and Adaptive Synchronization 30 Oct 2019 · 2 repositories · arXiv:1910.13598Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Lsh-sampling Breaks the Computation Chicken-and-egg Loop in Adaptive Stochastic Gradient Estimation 30 Oct 2019 · 0 repositories · arXiv:1910.14162
-
Online Stochastic Gradient Descent with Arbitrary Initialization Solves Non-smooth, Non-convex Phase Retrieval 28 Oct 2019 · 0 repositories · arXiv:1910.12837
-
A geometric interpretation of stochastic gradient descent using diffusion metrics 27 Oct 2019 · 0 repositories · arXiv:1910.12194
-
Asynchronous Decentralized SGD with Quantized and Local Updates 27 Oct 2019 · 0 repositories · arXiv:1910.12308
-
Sound Event Recognition in a Smart City Surveillance Context 27 Oct 2019 · 0 repositories · arXiv:1910.12369
-
Bias-Variance Tradeoff in a Sliding Window Implementation of the Stochastic Gradient Algorithm 25 Oct 2019 · 0 repositories · arXiv:1910.11868
-
The Practicality of Stochastic Optimization in Imaging Inverse Problems 22 Oct 2019 · 0 repositories · arXiv:1910.10100
-
Communication-Efficient Local Decentralized SGD Methods 21 Oct 2019 · 0 repositories · arXiv:1910.09126
-
Sparsification as a Remedy for Staleness in Distributed Asynchronous SGD 21 Oct 2019 · 0 repositories · arXiv:1910.09466
-
Error Lower Bounds of Constant Step-size Stochastic Gradient Descent 18 Oct 2019 · 0 repositories · arXiv:1910.08212
-
Improving the convergence of SGD through adaptive batch sizes 18 Oct 2019 · 0 repositories · arXiv:1910.08222
-
Interpreting Basis Path Set in Neural Networks 18 Oct 2019 · 0 repositories · arXiv:1910.09402
-
Robust Learning Rate Selection for Stochastic Optimization via Splitting Diagnostic 18 Oct 2019 · 0 repositories · arXiv:1910.08597
-
Why bigger is not always better: on finite and infinite neural networks 17 Oct 2019 · 0 repositories · arXiv:1910.08013
-
Derivative-Free Optimization of Neural Networks using Local Search 15 Oct 2019 · 1 repository
-
DP-MAC: The Differentially Private Method of Auxiliary Coordinates for Deep Learning 15 Oct 2019 · 1 repository · arXiv:1910.06924
-
Demon: Improved Neural Network Training with Momentum Decay 11 Oct 2019 · 2 repositories · arXiv:1910.04952
-
Distributed Learning of Deep Neural Networks using Independent Subnet Training 4 Oct 2019 · 2 repositories · arXiv:1910.02120
-
The Complexity of Finding Stationary Points with Stochastic Gradient Descent 4 Oct 2019 · 0 repositories · arXiv:1910.01845
-
Accelerating Deep Learning by Focusing on the Biggest Losers 2 Oct 2019 · 2 repositories · arXiv:1910.00762
-
Adaptive Activation Thresholding: Dynamic Routing Type Behavior for Interpretability in Convolutional Neural Networks 1 Oct 2019 · 0 repositories
-
How noise affects the Hessian spectrum in overparameterized neural networks 1 Oct 2019 · 0 repositories · arXiv:1910.00195
-
SlowMo: Improving Communication-Efficient Distributed SGD with Slow Momentum 1 Oct 2019 · 2 repositories · arXiv:1910.00643Syntology official: harvested for another paper · 0 ran · 3 unverified (of 3 harvested samples) · 1 pointer-only (licence)
-
Small Steps and Giant Leaps: Minimal Newton Solvers for Deep Learning 1 Oct 2019 · 1 repository
-
Distributed SGD Generalizes Well Under Asynchrony 29 Sep 2019 · 0 repositories · arXiv:1909.13391
-
A Mean-Field Theory for Kernel Alignment with Random Features in Generative and Discriminative Models 25 Sep 2019 · 0 repositories · arXiv:1909.11820
-
Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization 25 Sep 2019 · 2 repositories · arXiv:1909.11378
-
Gap Aware Mitigation of Gradient Staleness 24 Sep 2019 · 0 repositories · arXiv:1909.10802
-
Algorithm for Training Neural Networks on Resistive Device Arrays 17 Sep 2019 · 0 repositories · arXiv:1909.07908
-
Finite Depth and Width Corrections to the Neural Tangent Kernel 13 Sep 2019 · 0 repositories · arXiv:1909.05989
-
diffGrad: An Optimization Method for Convolutional Neural Networks 12 Sep 2019 · 1 repository · arXiv:1909.11015
-
The Error-Feedback Framework: Better Rates for SGD with Delayed Gradients and Compressed Communication 11 Sep 2019 · 0 repositories · arXiv:1909.05350
-
Tighter Theory for Local SGD on Identical and Heterogeneous Data 10 Sep 2019 · 0 repositories · arXiv:1909.04746
-
Byzantine-Resilient Stochastic Gradient Descent for Distributed Learning: A Lipschitz-Inspired Coordinate-wise Median Approach 10 Sep 2019 · 0 repositories · arXiv:1909.04532
-
A Stochastic Quasi-Newton Method with Nesterov's Accelerated Gradient 9 Sep 2019 · 0 repositories · arXiv:1909.03621
-
Communication-Censored Distributed Stochastic Gradient Descent 9 Sep 2019 · 1 repository · arXiv:1909.03631
-
Distributed Training of Embeddings using Graph Analytics 8 Sep 2019 · 0 repositories · arXiv:1909.03359
-
Decentralized Stochastic Gradient Tracking for Non-convex Empirical Risk Minimization 6 Sep 2019 · 0 repositories · arXiv:1909.02712
-
FreeAnchor: Learning to Match Anchors for Visual Object Detection 5 Sep 2019 · 4 repositories · arXiv:1909.02466Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Quasi-Newton Optimization Methods For Deep Learning Applications 4 Sep 2019 · 0 repositories · arXiv:1909.01994
-
A Concert-planning Tool for Independent Musicians by Machine Learning Models 29 Aug 2019 · 0 repositories · arXiv:1908.11200
-
Automatic and Simultaneous Adjustment of Learning Rate and Momentum for Stochastic Gradient Descent 20 Aug 2019 · 1 repository · arXiv:1908.07607
-
AdaCliP: Adaptive Clipping for Private SGD 20 Aug 2019 · 1 repository · arXiv:1908.07643
-
Multi Target Tracking by Learning from Generalized Graph Differences 19 Aug 2019 · 1 repository · arXiv:1908.06646
-
Towards Better Generalization: BP-SVRG in Training Deep Neural Networks 18 Aug 2019 · 0 repositories · arXiv:1908.06395
-
NUQSGD: Improved Communication Efficiency for Data-parallel SGD via Nonuniform Quantization 16 Aug 2019 · 1 repository · arXiv:1908.06077Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
On the Convergence of AdaBound and its Connection to SGD 13 Aug 2019 · 2 repositories · arXiv:1908.04457
-
Taming Unbalanced Training Workloads in Deep Learning with Partial Collective Operations 12 Aug 2019 · 0 repositories · arXiv:1908.04207
-
The HSIC Bottleneck: Deep Learning without Back-Propagation 5 Aug 2019 · 3 repositories · arXiv:1908.01580Syntology official (archive's flag): 2 ran · 4 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 12 unverified (of 16 harvested samples) · 1 pointer-only (licence)
-
How Good is SGD with Random Shuffling? 31 Jul 2019 · 0 repositories · arXiv:1908.00045
-
Deep Gradient Boosting -- Layer-wise Input Normalization of Neural Networks 29 Jul 2019 · 0 repositories · arXiv:1907.12608
-
DEAM: Adaptive Momentum with Discriminative Weight for Stochastic Optimization 25 Jul 2019 · 0 repositories · arXiv:1907.11307
-
Hessian based analysis of SGD for Deep Nets: Dynamics and Generalization 24 Jul 2019 · 0 repositories · arXiv:1907.10732
-
Mix and Match: An Optimistic Tree-Search Approach for Learning Models from Mixture Distributions 23 Jul 2019 · 1 repository · arXiv:1907.10154
-
Practical Newton-Type Distributed Learning using Gradient Based Approximations 22 Jul 2019 · 0 repositories · arXiv:1907.09562
-
Speeding Up Iterative Closest Point Using Stochastic Gradient Descent 22 Jul 2019 · 1 repository · arXiv:1907.09133
-
Post-synaptic potential regularization has potential 19 Jul 2019 · 1 repository · arXiv:1907.08544
-
SGD momentum optimizer with step estimation by online parabola model 16 Jul 2019 · 1 repository · arXiv:1907.07063
-
Amplifying Rényi Differential Privacy via Shuffling 11 Jul 2019 · 0 repositories · arXiv:1907.05156
-
A Highly Efficient Distributed Deep Learning System For Automatic Speech Recognition 10 Jul 2019 · 0 repositories · arXiv:1907.05701
-
Ordered SGD: A New Stochastic Optimization Framework for Empirical Risk Minimization 9 Jul 2019 · 2 repositories · arXiv:1907.04371Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples) · 2 pointer-only (licence)
-
Unified Optimal Analysis of the (Stochastic) Gradient Method 9 Jul 2019 · 0 repositories · arXiv:1907.04232
-
Stochastic Gradient and Langevin Processes 7 Jul 2019 · 0 repositories · arXiv:1907.03215
-
Next Generation Radiogenomics Sequencing for Prediction of EGFR and KRAS Mutation Status in NSCLC Patients Using Multimodal Imaging and Machine Learning Approaches 3 Jul 2019 · 0 repositories · arXiv:1907.02121
-
On Symmetry and Initialization for Neural Networks 1 Jul 2019 · 0 repositories · arXiv:1907.00560
-
Approximate matrix completion based on cavity method 29 Jun 2019 · 0 repositories · arXiv:1907.00138
-
DP-LSSGD: A Stochastic Optimization Method to Lift the Utility in Privacy-Preserving ERM 28 Jun 2019 · 1 repository · arXiv:1906.12056
-
Faster Distributed Deep Net Training: Computation and Communication Decoupled Stochastic Gradient Descent 28 Jun 2019 · 0 repositories · arXiv:1906.12043
-
Gradient Noise Convolution (GNC): Smoothing Loss Function for Distributed Large-Batch SGD 26 Jun 2019 · 0 repositories · arXiv:1906.10822
-
Learning Data Augmentation Strategies for Object Detection 26 Jun 2019 · 6 repositories · arXiv:1906.11172Syntology community repositories only · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples) · 3 pointer-only (licence)
-
First Exit Time Analysis of Stochastic Gradient Descent Under Heavy-Tailed Gradient Noise 21 Jun 2019 · 1 repository · arXiv:1906.09069
-
Data Cleansing for Models Trained with SGD 20 Jun 2019 · 1 repository · arXiv:1906.08473Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample)
-
Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup 18 Jun 2019 · 3 repositories · arXiv:1906.08632Syntology community repositories only · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 4 harvested samples)
-
On the Noisy Gradient Descent that Generalizes as SGD 18 Jun 2019 · 1 repository · arXiv:1906.07405
-
REMAP: Multi-layer entropy-guided pooling of dense CNN features for image retrieval 15 Jun 2019 · 0 repositories · arXiv:1906.06626
-
Stochastic Proximal AUC Maximization 14 Jun 2019 · 0 repositories · arXiv:1906.06053
-
Layered SGD: A Decentralized and Synchronous SGD Algorithm for Scalable Deep Neural Network Training 13 Jun 2019 · 0 repositories · arXiv:1906.05936
-
Training Neural Networks for and by Interpolation 13 Jun 2019 · 1 repository · arXiv:1906.05661Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples)
-
ADASS: Adaptive Sample Selection for Training Acceleration 11 Jun 2019 · 0 repositories · arXiv:1906.04819