Papers › Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods

Bayesian filtering unifies adaptive and non-adaptive neural network optimization methods

19 Jul 2018NeurIPS 2020 12arXiv:1807.07540archive 2025-07-28

Laurence Aitchison

We formulate the problem of neural network optimization as Bayesian filtering, where the observations are the backpropagated gradients. While neural network optimization has previously been studied using natural gradient methods which are closely related to Bayesian inference, they were unable to recover standard optimizers such as Adam and RMSprop with a root-mean-square gradient normalizer, instead getting a mean-square normalizer. To recover the root-mean-square normalizer, we find it necessary to account for the temporal dynamics of all the other parameters as they are geing optimized. The resulting optimizer, AdaBayes, adaptively transitions between SGD-like and Adam-like behaviour, automatically recovers AdamW, a state of the art variant of Adam with decoupled weight decay, and has generalisation performance competitive with SGD.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

LaurenceA/adabayes officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Bayesian Inference

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Methods

AdamAdamWRMSPropSGD

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections