Methods › General › Large Batch Optimization
Large Batch Optimization
Stochastic Optimization methods are used to optimize neural networks. We typically take a mini-batch of data, hence 'stochastic', and perform a type of gradient descent with this minibatch. Below you can find a continuously updating list of stochastic optimization algorithms.
Methods
All 8 methods in this collection, most-tagged first. Year is the archive's introduced_year; the archive stores 2000 when it has none, shown here as “–”. Papers counts distinct papers the archive tags with the method. Click a heading to sort.
| Adafactor | – | 733 |
| LAMB | – | 199 |
| AdaGrad | 2011 | 191 |
| LARS | – | 77 |
| 1-bit Adam | – | 40 |
| Nesterov Accelerated Gradient | 1983 | 34 |
| Distributed Shampoo | – | 5 |
| SLAMB Sparse Layer-wise Adaptive Moments optimizer for large Batch training | – | 1 |