Methods › General › Stochastic Optimization › SGD › Papers, page 9
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 9 of 21: papers 801 to 900 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Locally Asynchronous Stochastic Gradient Descent for Decentralised Deep Learning 24 Mar 2022 · 0 repositories · arXiv:2203.13085
-
A DNN Optimizer that Improves over AdaBelief by Suppression of the Adaptive Stepsize Range 24 Mar 2022 · 1 repository · arXiv:2203.13273
-
An Adaptive Gradient Method with Energy and Momentum 23 Mar 2022 · 1 repository · arXiv:2203.12191
-
ThingTalk: An Extensible, Executable Representation Language for Task-Oriented Dialogues 23 Mar 2022 · 1 repository · arXiv:2203.12751
-
Practical tradeoffs between memory, compute, and performance in learned optimizers 22 Mar 2022 · 1 repository · arXiv:2203.11860
-
Provable Constrained Stochastic Convex Optimization with XOR-Projected Gradient Descent 22 Mar 2022 · 0 repositories · arXiv:2203.11829
-
A Local Convergence Theory for the Stochastic Gradient Descent Method in Non-Convex Optimization With Non-isolated Local Minima 21 Mar 2022 · 0 repositories · arXiv:2203.10973
-
A new perspective on probabilistic image modeling 21 Mar 2022 · 0 repositories · arXiv:2203.11034
-
ImageNet Challenging Classification with the Raspberry Pi: An Incremental Local Stochastic Gradient Descent Algorithm 21 Mar 2022 · 0 repositories · arXiv:2203.11853
-
Fake News Detection Using Majority Voting Technique 18 Mar 2022 · 0 repositories · arXiv:2203.09936
-
Randomized Sharpness-Aware Training for Boosting Computational Efficiency in Deep Learning 18 Mar 2022 · 0 repositories · arXiv:2203.09962
-
The Role of Local Steps in Local SGD 14 Mar 2022 · 0 repositories · arXiv:2203.06798
-
Scaling the Wild: Decentralizing Hogwild!-style Shared-memory SGD 13 Mar 2022 · 1 repository · arXiv:2203.06638
-
Enhancing Adversarial Training with Second-Order Statistics of Weights 11 Mar 2022 · 1 repository · arXiv:2203.06020
-
Differentially Private Learning Needs Hidden State (Or Much Faster Convergence) 10 Mar 2022 · 0 repositories · arXiv:2203.05363
-
Risk Bounds of Multi-Pass SGD for Least Squares in the Interpolation Regime 7 Mar 2022 · 0 repositories · arXiv:2203.03159
-
What Did You Say? Task-Oriented Dialog Datasets Are Not Conversational!? 7 Mar 2022 · 0 repositories · arXiv:2203.03431
-
Towards Efficient and Scalable Sharpness-Aware Minimization 5 Mar 2022 · 4 repositories · arXiv:2203.02714Syntology 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample) · 1 pointer-only (licence)
-
Distributed Methods with Absolute Compression and Error Compensation 4 Mar 2022 · 0 repositories · arXiv:2203.02383
-
Benign Underfitting of Stochastic Gradient Descent 27 Feb 2022 · 0 repositories · arXiv:2202.13361
-
Explicit Regularization via Regularizer Mirror Descent 22 Feb 2022 · 0 repositories · arXiv:2202.10788
-
Personalized Federated Learning with Exact Stochastic Gradient Descent 20 Feb 2022 · 0 repositories · arXiv:2202.09848
-
On the Implicit Bias Towards Minimal Depth of Deep Neural Networks 18 Feb 2022 · 0 repositories · arXiv:2202.09028
-
Tackling benign nonconvexity with smoothing and stochastic gradients 18 Feb 2022 · 0 repositories · arXiv:2202.09052
-
Federated Stochastic Gradient Descent Begets Self-Induced Momentum 17 Feb 2022 · 0 repositories · arXiv:2202.08402
-
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks 17 Feb 2022 · 0 repositories · arXiv:2202.08658
-
Cost-Efficient Distributed Learning via Combinatorial Multi-Armed Bandits 16 Feb 2022 · 0 repositories · arXiv:2202.08302
-
Black-Box Generalization: Stability of Zeroth-Order Learning 14 Feb 2022 · 0 repositories · arXiv:2202.06880
-
Orthogonalising gradients to speed up neural network optimisation 14 Feb 2022 · 1 repository · arXiv:2202.07052
-
Escaping Saddle Points with Bias-Variance Reduced Local Perturbed SGD for Communication Efficient Nonconvex Distributed Learning 12 Feb 2022 · 0 repositories · arXiv:2202.06083
-
Maximizing Communication Efficiency for Large-scale Training via 0/1 Adam 12 Feb 2022 · 1 repository · arXiv:2202.06009
-
The Power of Adaptivity in SGD: Self-Tuning Step Sizes with Unbounded Gradients and Affine Variance 11 Feb 2022 · 0 repositories · arXiv:2202.05791
-
Robust Linear Regression for General Feature Distribution 4 Feb 2022 · 0 repositories · arXiv:2202.02080
-
SignSGD: Fault-Tolerance to Blind and Byzantine Adversaries 4 Feb 2022 · 3 repositories · arXiv:2202.02085
-
Characterizing & Finding Good Data Orderings for Fast Convergence of Sequential Gradient Methods 3 Feb 2022 · 0 repositories · arXiv:2202.01838
-
Fast Convex Optimization for Two-Layer ReLU Networks: Equivalent Model Classes and Cone Decompositions 2 Feb 2022 · 2 repositories · arXiv:2202.01331Syntology official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 6 with no instrument failure: 0 honoured, 0 violated, 6 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 7 harvested samples)
-
Robust Training of Neural Networks Using Scale Invariant Architectures 2 Feb 2022 · 0 repositories · arXiv:2202.00980
-
Phase diagram of Stochastic Gradient Descent in high-dimensional two-layer neural networks 1 Feb 2022 · 2 repositories · arXiv:2202.00293
-
Faster Convergence of Local SGD for Over-Parameterized Models 30 Jan 2022 · 0 repositories · arXiv:2201.12719
-
A Simple Guard for Learned Optimizers 28 Jan 2022 · 0 repositories · arXiv:2201.12426
-
DropNAS: Grouped Operation Dropout for Differentiable Architecture Search 27 Jan 2022 · 1 repository · arXiv:2201.11679
-
On the Convergence of mSGD and AdaGrad for Stochastic Optimization 26 Jan 2022 · 0 repositories · arXiv:2201.11204
-
On Uniform Boundedness Properties of SGD and its Momentum Variants 25 Jan 2022 · 0 repositories · arXiv:2201.10245
-
Description-Driven Task-Oriented Dialog Modeling 21 Jan 2022 · 1 repository · arXiv:2201.08904
-
Low-Pass Filtering SGD for Recovering Flat Optima in the Deep Learning Optimization Landscape 20 Jan 2022 · 1 repository · arXiv:2201.08025
-
AdaTerm: Adaptive T-Distribution Estimated Robust Moments for Noise-Robust Stochastic Gradient Optimization 18 Jan 2022 · 1 repository · arXiv:2201.06714
-
Unsupervised Slot Schema Induction for Task-oriented Dialog 16 Jan 2022 · 0 repositories
-
On generalization bounds for deep networks based on loss surface implicit regularization 12 Jan 2022 · 1 repository · arXiv:2201.04545
-
Partial Model Averaging in Federated Learning: Performance Guarantees and Benefits 11 Jan 2022 · 0 repositories · arXiv:2201.03789
-
Stability Based Generalization Bounds for Exponential Family Langevin Dynamics 9 Jan 2022 · 0 repositories · arXiv:2201.03064
-
The dynamics of representation learning in shallow, non-linear autoencoders 6 Jan 2022 · 1 repository · arXiv:2201.02115Syntology official: harvested, nothing ran · 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Stochastic regularized majorization-minimization with weakly convex and multi-convex surrogates 5 Jan 2022 · 1 repository · arXiv:2201.01652
-
Evaluation of Thermal Imaging on Embedded GPU Platforms for Application in Vehicular Assistance Systems 5 Jan 2022 · 0 repositories · arXiv:2201.01661
-
A Mixed-Integer Programming Approach to Training Dense Neural Networks 3 Jan 2022 · 0 repositories · arXiv:2201.00723
-
Stochastic Weight Averaging Revisited 3 Jan 2022 · 1 repository · arXiv:2201.00519
-
Accelerating Neural Network Optimization Through an Automated Control Theory Lens 1 Jan 2022 · 0 repositories
-
Distributed and Stochastic Optimization Methods with Gradient Compression and Local Steps 20 Dec 2021 · 0 repositories · arXiv:2112.10645
-
The effective noise of Stochastic Gradient Descent 20 Dec 2021 · 0 repositories · arXiv:2112.10852
-
Adaptively Customizing Activation Functions for Various Layers 17 Dec 2021 · 1 repository · arXiv:2112.09442
-
Non-Asymptotic Analysis of Online Multiplicative Stochastic Gradient Descent 14 Dec 2021 · 0 repositories · arXiv:2112.07110
-
Convergence proof for stochastic gradient descent in the training of deep neural networks with ReLU activation for constant target functions 13 Dec 2021 · 0 repositories · arXiv:2112.07369
-
Determinantal point processes based on orthogonal polynomials for sampling minibatches in SGD 11 Dec 2021 · 0 repositories · arXiv:2112.06007
-
Federated Two-stage Learning with Sign-based Voting 10 Dec 2021 · 0 repositories · arXiv:2112.05687
-
DR3: Value-Based Deep Reinforcement Learning Requires Explicit Regularization 9 Dec 2021 · 0 repositories · arXiv:2112.04716
-
FastSGD: A Fast Compressed SGD Framework for Distributed Machine Learning 8 Dec 2021 · 0 repositories · arXiv:2112.04291
-
Training Structured Neural Networks Through Manifold Identification and Variance Reduction 5 Dec 2021 · 2 repositories · arXiv:2112.02612Syntology official (archive's flag): 4 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 6 harvested samples) · 1 pointer-only (licence)
-
Loss Landscape Dependent Self-Adjusting Learning Rates in Decentralized Stochastic Gradient Descent 2 Dec 2021 · 0 repositories · arXiv:2112.01433
-
On Large Batch Training and Sharp Minima: A Fokker-Planck Perspective 2 Dec 2021 · 0 repositories · arXiv:2112.00987
-
Adaptive Proximal Gradient Methods for Structured Neural Networks 1 Dec 2021 · 0 repositories
-
An Improved Analysis and Rates for Variance Reduction under Without-replacement Sampling Orders 1 Dec 2021 · 0 repositories
-
Closing the Gap: Tighter Analysis of Alternating Stochastic Gradient Methods for Bilevel Problems 1 Dec 2021 · 0 repositories
-
Generalization Guarantee of SGD for Pairwise Learning 1 Dec 2021 · 0 repositories
-
Simple Stochastic and Online Gradient Descent Algorithms for Pairwise Learning 1 Dec 2021 · 0 repositories
-
The Implicit Bias of Minima Stability: A View from Function Space 1 Dec 2021 · 0 repositories
-
Towards Understanding Why Lookahead Generalizes Better Than SGD and Beyond 1 Dec 2021 · 1 repository
-
AutoDrop: Training Deep Learning Models with Automatic Learning Rate Drop 30 Nov 2021 · 0 repositories · arXiv:2111.15317
-
Randomized Stochastic Gradient Descent Ascent 25 Nov 2021 · 0 repositories · arXiv:2111.13162
-
Simple Stochastic and Online Gradient DescentAlgorithms for Pairwise Learning 23 Nov 2021 · 1 repository · arXiv:2111.12050Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Gaussian Process Inference Using Mini-batch Stochastic Gradient Descent: Convergence Guarantees and Empirical Benefits 19 Nov 2021 · 0 repositories · arXiv:2111.10461
-
An asynchronous distributed training algorithm based on Gossip 16 Nov 2021 · 0 repositories
-
Description-Driven Task-Oriented Dialog Modeling 16 Nov 2021 · 0 repositories
-
Dynamic Schema Graph Fusion Network for Multi-Domain Dialogue State Tracking 16 Nov 2021 · 0 repositories
-
Image-specific Convolutional Kernel Modulation for Single Image Super-resolution 16 Nov 2021 · 1 repository · arXiv:2111.08362
-
Improving Compositional Generalization with Self-Training for Data-to-Text Generation 16 Nov 2021 · 0 repositories
-
Online Meta Adaptation for Variable-Rate Learned Image Compression 16 Nov 2021 · 0 repositories · arXiv:2111.08256
-
Self-supervised Schema Induction for Task-oriented Dialog 16 Nov 2021 · 0 repositories
-
QK Iteration: A Self-Supervised Representation Learning Algorithm for Image Similarity 15 Nov 2021 · 0 repositories · arXiv:2111.07954
-
Stochastic Gradient Line Bayesian Optimization for Efficient Noise-Robust Optimization of Parameterized Quantum Circuits 15 Nov 2021 · 0 repositories · arXiv:2111.07952
-
The Three Stages of Learning Dynamics in High-Dimensional Kernel Methods 13 Nov 2021 · 0 repositories · arXiv:2111.07167
-
Stationary Behavior of Constant Stepsize SGD Type Algorithms: An Asymptotic Characterization 11 Nov 2021 · 0 repositories · arXiv:2111.06328
-
SGD Through the Lens of Kolmogorov Complexity 10 Nov 2021 · 0 repositories · arXiv:2111.05478
-
Learning to Rectify for Robust Learning with Noisy Labels 8 Nov 2021 · 1 repository · arXiv:2111.04239
-
Exponential escape efficiency of SGD from sharp minima in non-stationary regime 7 Nov 2021 · 1 repository · arXiv:2111.04004
-
AGGLIO: Global Optimization for Locally Convex Functions 6 Nov 2021 · 1 repository · arXiv:2111.03932
-
Sharp Bounds for Federated Averaging (Local SGD) and Continuous Perspective 5 Nov 2021 · 1 repository · arXiv:2111.03741
-
A Fast Parallel Tensor Decomposition with Optimal Stochastic Gradient Descent: an Application in Structural Damage Identification 4 Nov 2021 · 0 repositories · arXiv:2111.02632
-
Mean-field Analysis of Piecewise Linear Solutions for Wide ReLU Networks 3 Nov 2021 · 0 repositories · arXiv:2111.02278
-
Regularization by Misclassification in ReLU Neural Networks 3 Nov 2021 · 0 repositories · arXiv:2111.02154
-
Dropout in Training Neural Networks: Flatness of Solution and Noise Structure 1 Nov 2021 · 0 repositories · arXiv:2111.01022
-
Predicting Cancer Using Supervised Machine Learning: Mesothelioma 31 Oct 2021 · 0 repositories · arXiv:2111.01912