Methods › General › Stochastic Optimization
Stochastic Optimization
The archive attaches this collection's text per method and the copies differ: 2 distinct texts across 60 of the 65 methods here. All are shown, most-carried first (a tie goes to the text carrying Papers with Code's collection boilerplate, then to the longer text); no vote is taken between them.
Text 1, carried by 59 of 65 methods:
Stochastic Optimization methods are used to optimize neural networks. We typically take a mini-batch of data, hence 'stochastic', and perform a type of gradient descent with this minibatch. Below you can find a continuously updating list of stochastic optimization algorithms.
Text 2, carried by 1 of 65 methods:
This section contains a compilation of distributed methods for scaling deep learning to very large models. There are many different strategies for scaling training across multiple devices, including:
-
Data Parallel : for each node we use the same model parameters to do forward propagation, but we send a small batch of different data to each node, compute the gradient normally, and send it back to the main node. Once we have all the gradients, we calculate the weighted average and use this to update the model parameters.
-
Model Parallel : for each node we assign different layers to it. During forward propagation, we start in the node with the first layers, then move onto the next, and so on. Once forward propagation is done we calculate gradients for the last node, and update model parameters for that node. Then we backpropagate onto the penultimate node, update the parameters, and so on.
-
Additional methods including Hybrid Parallel, Auto Parallel, and Distributed Communication.
Image credit: Jordi Torres.
Methods
All 65 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| Adam | – | 24,390 |
| SGD Stochastic Gradient Descent | 1951 | 2,021 |
| ADOPT ADaptive gradient method with the OPTimal convergence rate | – | 831 |
| Adafactor | – | 733 |
| RMSProp | 2013 | 519 |
| MAS Mixing Adam and SGD | – | 207 |
| AdamW | – | 206 |
| Gravity | – | 202 |
| AdaGrad | 2011 | 191 |
| Deep Ensembles | – | 177 |
| SGD with Momentum | 1999 | 144 |
| FA Feedback Alignment | – | 124 |
| Local SGD | – | 69 |
| AO Artemisinin Optimization based on Malaria Therapy: Algorithm and Applications to Medical Image Segmentation | – | 66 |
| RAdam | – | 65 |
| Apollo Adaptive Parameter-wise Diagonal Quasi-Newton Method | – | 55 |
| DFA Direct Feedback Alignment | – | 55 |
| AMSGrad | – | 49 |
| Stochastic Weight Averaging | – | 45 |
| 1-bit Adam | – | 40 |
| Gradient Sparsification | – | 38 |
| PO Parrot optimizer: Algorithm and applications to medical problems | – | 37 |
| INFO INFO: An Efficient Optimization Algorithm based on Weighted Mean of Vectors | – | 36 |
| Nesterov Accelerated Gradient | 1983 | 34 |
| KP Kollen-Pollack Learning | – | 31 |
| ECO The Educational Competition Optimizer | – | 24 |
| Lookahead | – | 21 |
| SMA Slime Mould Algorithm | – | 21 |
| Adabelief | – | 19 |
| AdaDelta | – | 17 |
| Gradient Checkpointing | – | 14 |
| Forward gradient | – | 12 |
| HGS Hunger Games Search | – | 12 |
| AdaBound | – | 11 |
| AdaMax | – | 9 |
| NADAM | 2015 | 7 |
| SM3 | – | 7 |
| AdaHessian ADAHESSIAN | – | 6 |
| NT-ASGD Non-monotonically Triggered ASGD | – | 6 |
| Distributed Shampoo | – | 5 |
| AdaShift | – | 4 |
| Polyak Averaging | 1991 | 3 |
| PowerSGD | – | 3 |
| AdaFisher Adaptive Second Order Optimization via Fisher Information | – | 2 |
| AdaSmooth Adaptive Smooth Optimizer | – | 2 |
| Adam-mini Adaptive Moment Estimation - Mini | – | 2 |
| MPSO Motion-Encoded Particle Swarm Optimization | – | 2 |
| QHM | – | 2 |
| SGDW | – | 2 |
| AMSBound | – | 1 |
| ATMO AdapTive Meta Optimizer | – | 1 |
| AdaMod | – | 1 |
| AdaSqrt | – | 1 |
| AggMo | – | 1 |
| DSPT double-stage parameter tuning | – | 1 |
| Demon ADAM | – | 1 |
| Demon CM | – | 1 |
| FATA FATA: An Efficient Optimization Method based on Geophysics | – | 1 |
| MADGRAD Momentumized, adaptive, dual averaged gradient | – | 1 |
| Powerpropagation | – | 1 |
| QHAdam | – | 1 |
| SRMM Stochastic Regularized Majorization-Minimization | – | 1 |
| YellowFin | – | 1 |
| FASFA FASFA: A Novel Next-Generation Backpropagation Optimizer | – | 0 |
| PLO Polar Lights Optimizer | – | 0 |