{"url":"/method/natural-gradient-descent","slug":"natural-gradient-descent","name":"Natural Gradient Descent","full_name":"Natural Gradient Descent","full_name_withheld":false,"description_markdown":"**Natural Gradient Descent** is an approximate second-order optimisation method. It has an interpretation as optimizing over a Riemannian manifold using an intrinsic distance metric, which implies the updates are invariant to transformations such as whitening. By using the positive semi-definite (PSD) Gauss-Newton matrix to approximate the (possibly negative definite) Hessian, NGD can often work better than exact second-order methods.\r\n\r\nGiven the gradient of $z$, $g = \\frac{\\delta{f}\\left(z\\right)}{\\delta{z}}$, NGD computes the update as:\r\n\r\n$$\\Delta{z} = \\alpha{F}^{−1}g$$\r\n\r\nwhere the Fisher information matrix $F$ is defined as:\r\n\r\n$$ F = \\mathbb{E}\\_{p\\left(t\\mid{z}\\right)}\\left[\\nabla\\ln{p}\\left(t\\mid{z}\\right)\\nabla\\ln{p}\\left(t\\mid{z}\\right)^{T}\\right] $$\r\n\r\nThe log-likelihood function $\\ln{p}\\left(t\\mid{z}\\right)$ typically corresponds to commonly used error functions such as the cross entropy loss.\r\n\r\nSource: [LOGAN](https://paperswithcode.com/method/logan)\r\n\r\nImage: [Fast Convergence of Natural Gradient Descent for Overparameterized Neural Networks\r\n](https://arxiv.org/abs/1905.10961)","description_state":"present","introduced_year":1998,"introduced_by":{"title":null,"paper":null,"first_author":null,"n_authors":0,"url_abs":null,"archive_paper_url":null},"source":{"url":null,"title":null,"url_on_a_paper_host":false},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Optimization","url":"/methods/category/optimization","pwc_aliases":[]}],"n_papers_tagged":68,"archive_num_papers":null,"papers_newest_first":[{"paper":null,"title":"Improving Energy Natural Gradient Descent through Woodbury, Momentum, and Randomization","date":"2025-05-17","arxiv_id":"2505.12149","n_code_links":0,"syntology":null},{"paper":null,"title":"Convergence Properties of Natural Gradient Descent for Minimizing KL Divergence","date":"2025-04-27","arxiv_id":"2504.19259","n_code_links":0,"syntology":null},{"paper":"/paper/a-mean-teacher-algorithm-for-unlearning-of","title":"A mean teacher algorithm for unlearning of language models","date":"2025-04-18","arxiv_id":"2504.13388","n_code_links":1,"syntology":null},{"paper":null,"title":"Natural Gradient Descent for Control","date":"2025-03-08","arxiv_id":"2503.06070","n_code_links":0,"syntology":null},{"paper":null,"title":"Guiding Time-Varying Generative Models with Natural Gradients on Exponential Family Manifold","date":"2025-02-11","arxiv_id":"2502.07650","n_code_links":0,"syntology":null},{"paper":"/paper/high-accuracy-sampling-from-constrained","title":"High-accuracy sampling from constrained spaces with the Metropolis-adjusted Preconditioned Langevin Algorithm","date":"2024-12-24","arxiv_id":"2412.18701","n_code_links":1,"syntology":{"ran":0,"of":3,"unverified":3,"pointer_only":0}},{"paper":"/paper/reconstructing-deep-neural-networks","title":"Reconstructing Deep Neural Networks: Unleashing the Optimization Potential of Natural Gradient Descent","date":"2024-12-10","arxiv_id":"2412.07441","n_code_links":1,"syntology":null},{"paper":null,"title":"Natural gradient and parameter estimation for quantum Boltzmann machines","date":"2024-10-31","arxiv_id":"2410.24058","n_code_links":0,"syntology":null},{"paper":"/paper/a-prescriptive-theory-for-brain-like","title":"Brain-like variational inference","date":"2024-10-25","arxiv_id":"2410.19315","n_code_links":0,"syntology":{"ran":0,"of":6,"unverified":6,"pointer_only":0}},{"paper":null,"title":"Score-Based Variational Inference for Inverse Problems","date":"2024-10-08","arxiv_id":"2410.05646","n_code_links":0,"syntology":null},{"paper":null,"title":"Is All Learning (Natural) Gradient Descent?","date":"2024-09-24","arxiv_id":"2409.16422","n_code_links":0,"syntology":null},{"paper":"/paper/ngd-converges-to-less-degenerate-solutions","title":"NGD converges to less degenerate solutions than SGD","date":"2024-09-07","arxiv_id":"2409.04913","n_code_links":1,"syntology":null},{"paper":null,"title":"Decentralised Variational Inference Frameworks for Multi-object Tracking on Sensor Networks: Additional Notes","date":"2024-08-24","arxiv_id":"2408.13689","n_code_links":0,"syntology":null},{"paper":null,"title":"Convergence Analysis of Natural Gradient Descent for Over-parameterized Physics-Informed Neural Networks","date":"2024-08-01","arxiv_id":"2408.00573","n_code_links":0,"syntology":null},{"paper":"/paper/quantum-natural-stochastic-pairwise","title":"Quantum Natural Stochastic Pairwise Coordinate Descent","date":"2024-07-18","arxiv_id":"2407.13858","n_code_links":1,"syntology":null},{"paper":"/paper/correlations-are-ruining-your-gradient","title":"Correlations Are Ruining Your Gradient Descent","date":"2024-07-15","arxiv_id":"2407.10780","n_code_links":1,"syntology":null},{"paper":null,"title":"Faster Machine Unlearning via Natural Gradient Descent","date":"2024-07-11","arxiv_id":"2407.08169","n_code_links":0,"syntology":null},{"paper":null,"title":"An Improved Empirical Fisher Approximation for Natural Gradient Descent","date":"2024-06-10","arxiv_id":"2406.06420","n_code_links":0,"syntology":null},{"paper":"/paper/bayesian-online-natural-gradient-bong","title":"Bayesian Online Natural Gradient (BONG)","date":"2024-05-30","arxiv_id":"2405.19681","n_code_links":1,"syntology":{"ran":8,"of":9,"unverified":1,"pointer_only":0}},{"paper":null,"title":"Thermodynamic Natural Gradient Descent","date":"2024-05-22","arxiv_id":"2405.13817","n_code_links":0,"syntology":null},{"paper":"/paper/fadam-adam-is-a-natural-gradient-optimizer","title":"FAdam: Adam is a natural gradient optimizer using diagonal empirical Fisher information","date":"2024-05-21","arxiv_id":"2405.12807","n_code_links":1,"syntology":null},{"paper":null,"title":"Sequential-in-time training of nonlinear parametrizations for solving time-dependent partial differential equations","date":"2024-04-01","arxiv_id":"2404.01145","n_code_links":0,"syntology":null},{"paper":null,"title":"Inverse-Free Fast Natural Gradient Descent Method for Deep Learning","date":"2024-03-06","arxiv_id":"2403.03473","n_code_links":0,"syntology":null},{"paper":null,"title":"SOFIM: Stochastic Optimization Using Regularized Fisher Information Matrix","date":"2024-03-05","arxiv_id":"2403.02833","n_code_links":0,"syntology":null},{"paper":"/paper/a-kaczmarz-inspired-approach-to-accelerate","title":"A Kaczmarz-inspired approach to accelerate the optimization of neural network wavefunctions","date":"2024-01-18","arxiv_id":"2401.10190","n_code_links":1,"syntology":null},{"paper":"/paper/structured-inverse-free-natural-gradient","title":"Structured Inverse-Free Natural Gradient: Memory-Efficient & Numerically-Stable KFAC","date":"2023-12-09","arxiv_id":"2312.05705","n_code_links":2,"syntology":{"ran":1,"of":1,"unverified":0,"pointer_only":1}},{"paper":null,"title":"Efficient Numerical Algorithm for Large-Scale Damped Natural Gradient Descent","date":"2023-10-26","arxiv_id":"2310.17556","n_code_links":0,"syntology":null},{"paper":null,"title":"Modify Training Directions in Function Space to Reduce Generalization Error","date":"2023-07-25","arxiv_id":"2307.13290","n_code_links":0,"syntology":null},{"paper":"/paper/analysis-and-comparison-of-two-level-kfac","title":"Analysis and Comparison of Two-Level KFAC Methods for Training Deep Neural Networks","date":"2023-03-31","arxiv_id":"2303.18083","n_code_links":1,"syntology":null},{"paper":null,"title":"Decentralized Riemannian natural gradient methods with Kronecker-product approximations","date":"2023-03-16","arxiv_id":"2303.09611","n_code_links":0,"syntology":null}],"papers_shown":30,"tasks":[{"task":"/task/variational-inference","name":"Variational Inference","papers":8},{"task":"/task/image-classification","name":"Image Classification","papers":7},{"task":"/task/image-classification","name":"image-classification","papers":7},{"task":"/task/second-order-methods","name":"Second-order methods","papers":5},{"task":"/task/computational-efficiency","name":"Computational Efficiency","papers":4},{"task":"/task/deep-learning","name":"Deep Learning","papers":3},{"task":"/task/quantum-machine-learning","name":"Quantum Machine Learning","papers":3},{"task":"/task/stochastic-optimization","name":"Stochastic Optimization","papers":3},{"task":"/task/variational-monte-carlo","name":"Variational Monte Carlo","papers":3},{"task":"/task/machine-learning","name":"BIG-bench Machine Learning","papers":2},{"task":"/task/bayesian-inference","name":"Bayesian Inference","papers":2},{"task":"/task/bias-detection","name":"Bias Detection","papers":2},{"task":"/task/clustering","name":"Clustering","papers":2},{"task":"/task/federated-learning","name":"Federated Learning","papers":2},{"task":"/task/image-reconstruction","name":"Image Reconstruction","papers":2},{"task":"/task/language-modeling","name":"Language Modeling","papers":2},{"task":"/task/language-modelling","name":"Language Modelling","papers":2},{"task":"/task/regression-1","name":"regression","papers":2},{"task":"/task/3d-reconstruction","name":"3D Reconstruction","papers":1},{"task":"/task/adversarial-attack","name":"Adversarial Attack","papers":1}],"tasks_shown":20,"n_tasks":51,"usage_by_year":[{"year":"2019","papers":5},{"year":"2020","papers":9},{"year":"2021","papers":12},{"year":"2022","papers":9},{"year":"2023","papers":8},{"year":"2024","papers":20},{"year":"2025","papers":5}],"row_source":"embedded","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/natural-gradient-descent"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}