Methods › General › Large Batch Optimization › Distributed Shampoo
Distributed Shampoo
Introduced by Rohan Anil et al. in Towards Practical Second Order Optimization for Deep Learning
archive 2025-07-28 Description, source and code snippet are the archive's method entry.
A scalable second order optimization algorithm for deep learning.
Optimization in machine learning, both theoretical and applied, is presently dominated by first-order gradient methods such as stochastic gradient descent. Second-order optimization methods, that involve second derivatives and/or second order statistics of the data, are far less prevalent despite strong theoretical properties, due to their prohibitive computation, memory and communication costs. In an attempt to bridge this gap between theoretical and practical optimization, we present a scalable implementation of a second-order preconditioned method (concretely, a variant of full-matrix Adagrad), that along with several critical algorithmic and numerical improvements, provides significant convergence and wall-clock time improvements compared to conventional first-order methods on state-of-the-art deep models. Our novel design effectively utilizes the prevalent heterogeneous hardware architecture for training deep models, consisting of a multicore CPU coupled with multiple accelerator units. We demonstrate superior performance compared to state-of-the-art on very large learning tasks such as machine translation with Transformers, language modeling with BERT, click-through rate prediction on Criteo, and image classification on ImageNet with ResNet-50.
Papers archive 2025-07-28
5 shown of 5, newest first. Repository counts are the archive's code-links table. A Syntology line states what Syntology ran from that paper's harvested code; it is per sample and not a correctness claim.
-
Accelerating Neural Network Training: An Analysis of the AlgoPerf Competition 20 Feb 2025 · 2 repositories · arXiv:2502.15015Syntology ran 0 of 19 samples · 19 unverified
-
Knowledge distillation: A good teacher is patient and consistent 9 Jun 2021 · 10 repositories · arXiv:2106.05237
-
Tensor Normal Training for Deep Learning Models 5 Jun 2021 · 1 repository · arXiv:2106.02925Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)
-
Towards Practical Second Order Optimization for Deep Learning 1 Jan 2021 · 0 repositories
-
Scalable Second Order Optimization for Deep Learning 20 Feb 2020 · 2 repositories · arXiv:2002.09018
Tasks archive 2025-07-28
12 tasks the archive attaches to papers tagged with this method, by distinct papers. A task without a page in the catalog is plain text.
Usage over time archive 2025-07-28
Components: the archive holds no method-to-method composition, so PwC's Components table cannot be rebuilt; the Papers list carries no Results column for the same reason (the archive does not join its leaderboard rows to method tags).
Categories archive 2025-07-28
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections