Methods › General › Stochastic Optimization › SGD › Papers, page 11
Stochastic Gradient Descent
SGD
Papers archive 2025-07-28
archive papers tagged: 2,021 · with a code link: 591 · where Syntology ran a sample: 192 (161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (192 of 2,021 tagged: 161 with a run with no instrument failure, 31 where every run was a failure of Syntology's instrument)
Page 11 of 21: papers 1,001 to 1,100 of 2,021, newest first by the archive's date (ties by slug), in archive order.
Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code, as “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the instrument figure counts failures of Syntology's instrument, not of the code. It is per sample and not a correctness claim. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from.
-
Activated Gradients for Deep Neural Networks 9 Jul 2021 · 2 repositories · arXiv:2107.04228
-
REX: Revisiting Budgeted Training with an Improved Schedule 9 Jul 2021 · 1 repository · arXiv:2107.04197
-
M-FAC: Efficient Matrix-Free Approximations of Second-Order Information 7 Jul 2021 · 2 repositories · arXiv:2107.03356Syntology official (archive's flag): 4 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 4 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 7 unverified (of 14 harvested samples)
-
AdaL: Adaptive Gradient Transformation Contributes to Convergences and Generalizations 4 Jul 2021 · 0 repositories · arXiv:2107.01525
-
Exact Backpropagation in Binary Weighted Networks with Group Weight Transformations 3 Jul 2021 · 1 repository · arXiv:2107.01400
-
ResIST: Layer-Wise Decomposition of ResNets for Distributed Training 2 Jul 2021 · 0 repositories · arXiv:2107.00961
-
High-probability Bounds for Non-Convex Stochastic Optimization with Heavy Tails 28 Jun 2021 · 0 repositories · arXiv:2106.14343
-
The Convergence Rate of SGD's Final Iterate: Analysis on Dimension Dependence 28 Jun 2021 · 0 repositories · arXiv:2106.14588
-
LNS-Madam: Low-Precision Training in Logarithmic Number System using Multiplicative Weight Update 26 Jun 2021 · 0 repositories · arXiv:2106.13914
-
Implicit Gradient Alignment in Distributed and Federated Learning 25 Jun 2021 · 0 repositories · arXiv:2106.13897
-
Private Adaptive Gradient Methods for Convex Optimization 25 Jun 2021 · 0 repositories · arXiv:2106.13756
-
Tighter Analysis of Alternating Stochastic Gradient Method for Stochastic Nested Problems 25 Jun 2021 · 0 repositories · arXiv:2106.13781
-
Understanding Clipping for Federated Learning: Convergence and Client-Level Differential Privacy 25 Jun 2021 · 0 repositories · arXiv:2106.13673
-
Numerical influence of ReLU'(0) on backpropagation 23 Jun 2021 · 1 repository · arXiv:2106.12915
-
Stochastic Polyak Stepsize with a Moving Target 22 Jun 2021 · 0 repositories · arXiv:2106.11851
-
FedCM: Federated Learning with Client-level Momentum 21 Jun 2021 · 2 repositories · arXiv:2106.10874
-
How Do Adam and Training Strategies Help BNNs Optimization? 21 Jun 2021 · 0 repositories · arXiv:2106.11309
-
Open-set Label Noise Can Improve Robustness Against Inherent Label Noise 21 Jun 2021 · 4 repositories · arXiv:2106.10891Syntology official (archive's flag): 9 ran · 17 ran (of which 5 constructed an object rather than computing a result; 10 with no instrument failure: 1 honoured, 1 violated, 8 with no contract checked; 7 where Syntology's instrument failed) · 12 unverified (of 29 harvested samples) · 29 pointer-only (licence)
-
Multirate Training of Neural Networks 20 Jun 2021 · 4 repositories · arXiv:2106.10771Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
STEM: A Stochastic Two-Sided Momentum Algorithm Achieving Near-Optimal Sample and Communication Complexities for Federated Learning 19 Jun 2021 · 0 repositories · arXiv:2106.10435
-
Variational Prototype Learning for Deep Face Recognition 19 Jun 2021 · 0 repositories
-
Large Scale Private Learning via Low-rank Reparametrization 17 Jun 2021 · 1 repository · arXiv:2106.09352Syntology official (archive's flag): 4 ran · 4 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 3 where Syntology's instrument failed) · 4 unverified (of 8 harvested samples) · 8 pointer-only (licence)
-
Private Federated Learning Without a Trusted Server: Optimal Algorithms for Convex Losses 17 Jun 2021 · 2 repositories · arXiv:2106.09779Syntology official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 3 honoured, 1 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 4 harvested samples)
-
Effective Evaluation of Deep Active Learning on Image Classification Tasks 16 Jun 2021 · 0 repositories · arXiv:2106.15324
-
Exponential Error Convergence in Data Classification with Optimized Random Features: Acceleration by Quantum Machine Learning 16 Jun 2021 · 0 repositories · arXiv:2106.09028
-
Robust Training in High Dimensions via Block Coordinate Geometric Median Descent 16 Jun 2021 · 2 repositories · arXiv:2106.08882Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (of 3 harvested samples)
-
Masked Training of Neural Networks with Partial Gradients 16 Jun 2021 · 0 repositories · arXiv:2106.08895
-
Quantum-inspired event reconstruction with Tensor Networks: Matrix Product States 15 Jun 2021 · 1 repository · arXiv:2106.08334
-
Revisiting Model Stitching to Compare Neural Representations 14 Jun 2021 · 0 repositories · arXiv:2106.07682
-
A decreasing scaling transition scheme from Adam to SGD 12 Jun 2021 · 2 repositories · arXiv:2106.06749
-
Random Shuffling Beats SGD Only After Many Epochs on Ill-Conditioned Problems 12 Jun 2021 · 1 repository · arXiv:2106.06880
-
Label Noise SGD Provably Prefers Flat Global Minimizers 11 Jun 2021 · 0 repositories · arXiv:2106.06530
-
Communication-efficient SGD: From Local SGD to One-Shot Averaging 9 Jun 2021 · 0 repositories · arXiv:2106.04759
-
Fractal Structure and Generalization Properties of Stochastic Optimization Algorithms 9 Jun 2021 · 0 repositories · arXiv:2106.04881
-
Batch Normalization Orthogonalizes Representations in Deep Random Networks 7 Jun 2021 · 1 repository · arXiv:2106.03970
-
Dynamics of Stochastic Momentum Methods on Large-scale, Quadratic Models 7 Jun 2021 · 0 repositories · arXiv:2106.03696
-
Heavy Tails in SGD and Compressibility of Overparametrized Neural Networks 7 Jun 2021 · 1 repository · arXiv:2106.03795
-
Vanishing Curvature and the Power of Adaptive Methods in Randomly Initialized Deep Networks 7 Jun 2021 · 0 repositories · arXiv:2106.03763
-
Fast and Robust Online Inference with Stochastic Gradient Descent via Random Scaling 6 Jun 2021 · 1 repository · arXiv:2106.03156
-
Bandwidth-based Step-Sizes for Non-Convex Stochastic Optimization 5 Jun 2021 · 0 repositories · arXiv:2106.02888
-
Escaping Saddle Points Faster with Stochastic Momentum 5 Jun 2021 · 0 repositories · arXiv:2106.02985
-
Debiasing a First-order Heuristic for Approximate Bi-level Optimization 4 Jun 2021 · 1 repository · arXiv:2106.02487Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 3 harvested samples)
-
Learning Curves for SGD on Structured Features 4 Jun 2021 · 1 repository · arXiv:2106.02713
-
SpreadGNN: Serverless Multi-task Federated Learning for Graph Neural Networks 4 Jun 2021 · 1 repository · arXiv:2106.02743
-
Stochastic gradient descent with noise of machine learning type. Part II: Continuous time analysis 4 Jun 2021 · 0 repositories · arXiv:2106.02588
-
Continual Learning in Deep Networks: an Analysis of the Last Layer 3 Jun 2021 · 0 repositories · arXiv:2106.01834
-
LRTuner: A Learning Rate Tuner for Deep Neural Networks 30 May 2021 · 2 repositories · arXiv:2105.14526
-
On Linear Stability of SGD and Input-Smoothness of Neural Networks 27 May 2021 · 1 repository · arXiv:2105.13462Syntology official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 2 honoured, 0 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
Training With Data Dependent Dynamic Learning Rates 27 May 2021 · 0 repositories · arXiv:2105.13464
-
Using Early-Learning Regularization to Classify Real-World Noisy Data 27 May 2021 · 0 repositories · arXiv:2105.13244
-
Near-optimal Offline and Streaming Algorithms for Learning Non-Linear Dynamical Systems 24 May 2021 · 0 repositories · arXiv:2105.11558
-
GOALS: Gradient-Only Approximations for Line Searches Towards Robust and Consistent Training of Deep Neural Networks 23 May 2021 · 0 repositories · arXiv:2105.10915
-
AngularGrad: A New Optimization Technique for Angular Convergence of Convolutional Neural Networks 21 May 2021 · 3 repositories · arXiv:2105.10190
-
Escaping Saddle Points with Compressed SGD 21 May 2021 · 0 repositories · arXiv:2105.10090
-
Last iterate convergence of SGD for Least-Squares in the Interpolation regime. 21 May 2021 · 0 repositories
-
Numerical influence of ReLU’(0) on backpropagation 21 May 2021 · 0 repositories
-
On the Generalization of Neural Networks Trained with SGD: Information-Theoretical Bounds and Implications 21 May 2021 · 0 repositories
-
Online Statistical Inference for Parameters Estimation with Linear-Equality Constraints 21 May 2021 · 0 repositories · arXiv:2105.10315
-
Privacy Amplification Via Bernoulli Sampling 21 May 2021 · 0 repositories · arXiv:2105.10594
-
Properties of the After Kernel 21 May 2021 · 0 repositories · arXiv:2105.10585
-
Power-law escape rate of SGD 20 May 2021 · 0 repositories · arXiv:2105.09557
-
Towards Quantized Model Parallelism for Graph-Augmented MLPs Based on Gradient-Free ADMM Framework 20 May 2021 · 1 repository · arXiv:2105.09837
-
Accelerating Gossip SGD with Periodic Global Averaging 19 May 2021 · 0 repositories · arXiv:2105.09080
-
Removing Data Heterogeneity Influence Enhances Network Topology Dependence of Decentralized SGD 17 May 2021 · 0 repositories · arXiv:2105.08023
-
SGD-QA: Fast Schema-Guided Dialogue State Tracking for Unseen Services 17 May 2021 · 0 repositories · arXiv:2105.08049
-
Drill the Cork of Information Bottleneck by Inputting the Most Important Data 15 May 2021 · 0 repositories · arXiv:2105.07181
-
Why Does Multi-Epoch Training Help? 13 May 2021 · 0 repositories · arXiv:2105.06015
-
Evading the Simplicity Bias: Training a Diverse Set of Models Discovers Solutions with Superior OOD Generalization 12 May 2021 · 1 repository · arXiv:2105.05612
-
Tensor Programs IIb: Architectural Universality of Neural Tangent Kernel Training Dynamics 8 May 2021 · 0 repositories · arXiv:2105.03703
-
Understanding Short-Range Memory Effects in Deep Neural Networks 5 May 2021 · 0 repositories · arXiv:2105.02062
-
Information Complexity and Generalization Bounds 4 May 2021 · 0 repositories · arXiv:2105.01747
-
Stochastic gradient descent with noise of machine learning type. Part I: Discrete time analysis 4 May 2021 · 0 repositories · arXiv:2105.01650
-
A Perceptual Distortion Reduction Framework: Towards Generating Adversarial Examples with High Perceptual Quality and Attack Success Rate 1 May 2021 · 0 repositories · arXiv:2105.00278
-
One-pass Stochastic Gradient Descent in Overparametrized Two-layer Neural Networks 1 May 2021 · 0 repositories · arXiv:2105.00262
-
Studying the Consistency and Composability of Lottery Ticket Pruning Masks 30 Apr 2021 · 0 repositories · arXiv:2104.14753
-
NUQSGD: Provably Communication-efficient Data-parallel SGD via Nonuniform Quantization 28 Apr 2021 · 0 repositories · arXiv:2104.13818
-
Improved Analysis and Rates for Variance Reduction under Without-replacement Sampling Orders 25 Apr 2021 · 0 repositories · arXiv:2104.12112
-
DecentLaM: Decentralized Momentum SGD for Large-batch Deep Training 24 Apr 2021 · 1 repository · arXiv:2104.11981
-
Partitioning sparse deep neural networks for scalable training and inference 23 Apr 2021 · 0 repositories · arXiv:2104.11805
-
MetricOpt: Learning to Optimize Black-Box Evaluation Metrics 21 Apr 2021 · 0 repositories · arXiv:2104.10631
-
Random Reshuffling with Variance Reduction: New Analysis and Better Rates 19 Apr 2021 · 0 repositories · arXiv:2104.09342
-
A Method to Reveal Speaker Identity in Distributed ASR Training, and How to Counter It 15 Apr 2021 · 1 repository · arXiv:2104.07815
-
D-Cliques: Compensating for Data Heterogeneity with Topology in Decentralized Federated Learning 15 Apr 2021 · 0 repositories · arXiv:2104.07365
-
1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB's Convergence Speed 13 Apr 2021 · 1 repository · arXiv:2104.06069
-
Sample-based and Feature-based Federated Learning for Unconstrained and Constrained Nonconvex Optimization via Mini-batch SSCA 13 Apr 2021 · 1 repository · arXiv:2104.06011
-
BERT-based Chinese Text Classification for Emergency Domain with a Novel Loss Function 9 Apr 2021 · 0 repositories · arXiv:2104.04197
-
A proof of convergence for stochastic gradient descent in the training of artificial neural networks with ReLU activation for constant target functions 1 Apr 2021 · 0 repositories · arXiv:2104.00277
-
Empirically explaining SGD from a line search perspective 31 Mar 2021 · 1 repository · arXiv:2103.17132
-
Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to Improve Generalization 31 Mar 2021 · 1 repository · arXiv:2103.17182Syntology official (archive's flag): 1 ran · 1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result (of 1 harvested sample)
-
Research of Damped Newton Stochastic Gradient Descent Method for Neural Network Training 31 Mar 2021 · 0 repositories · arXiv:2103.16764
-
Exploiting Adam-like Optimization Algorithms to Improve the Performance of Convolutional Neural Networks 26 Mar 2021 · 0 repositories · arXiv:2103.14689
-
Adaptive Importance Sampling for Finite-Sum Optimization and Sampling with Decreasing Step-Sizes 23 Mar 2021 · 0 repositories · arXiv:2103.12243
-
Benign Overfitting of Constant-Stepsize SGD for Linear Regression 23 Mar 2021 · 0 repositories · arXiv:2103.12692
-
Stochastic Reweighted Gradient Descent 23 Mar 2021 · 0 repositories · arXiv:2103.12293
-
ProgressiveSpinalNet architecture for FC layers 21 Mar 2021 · 1 repository · arXiv:2103.11373
-
DataLens: Scalable Privacy Preserving Training via Gradient Compression and Aggregation 20 Mar 2021 · 2 repositories · arXiv:2103.11109
-
A deep learning theory for neural networks grounded in physics 18 Mar 2021 · 0 repositories · arXiv:2103.09985
-
Distributed Deep Learning Using Volunteer Computing-Like Paradigm 16 Mar 2021 · 0 repositories · arXiv:2103.08894
-
Hebbian Semi-Supervised Learning in a Sample Efficiency Setting 16 Mar 2021 · 0 repositories · arXiv:2103.09002
-
Repurposing Pretrained Models for Robust Out-of-domain Few-Shot Learning 16 Mar 2021 · 1 repository · arXiv:2103.09027Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)