{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-kronecker-factored-approximate-fisher","title":"A Kronecker-factored approximate Fisher matrix for convolution layers","arxiv_id":"1602.01407","date":"2016-02-03","proceeding":null,"authors":["Roger Grosse","James Martens"],"abstract":"Second-order optimization methods such as natural gradient descent have the\npotential to speed up training of neural networks by correcting for the\ncurvature of the loss function. Unfortunately, the exact natural gradient is\nimpractical to compute for large models, and most approximations either require\nan expensive iterative procedure or make crude approximations to the curvature.\nWe present Kronecker Factors for Convolution (KFC), a tractable approximation\nto the Fisher matrix for convolutional networks based on a structured\nprobabilistic model for the distribution over backpropagated derivatives.\nSimilarly to the recently proposed Kronecker-Factored Approximate Curvature\n(K-FAC), each block of the approximate Fisher matrix decomposes as the\nKronecker product of small matrices, allowing for efficient inversion. KFC\ncaptures important curvature information while still yielding comparably\nefficient updates to stochastic gradient descent (SGD). We show that the\nupdates are invariant to commonly used reparameterizations, such as centering\nof the activations. In our experiments, approximate natural gradient descent\nwith KFC was able to train convolutional networks several times faster than\ncarefully tuned SGD. Furthermore, it was able to train the networks in 10-20\ntimes fewer iterations than SGD, suggesting its potential applicability in a\ndistributed setting.","url_abs":"http://arxiv.org/abs/1602.01407v2","url_pdf":"http://arxiv.org/pdf/1602.01407v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-kronecker-factored-approximate-fisher","repo_url":"https://github.com/Thrandis/EKFAC-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"a-kronecker-factored-approximate-fisher","repo_url":"https://github.com/lzhangbv/eva","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"a-kronecker-factored-approximate-fisher","repo_url":"https://github.com/n-gao/pytorch-kfac","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"a-kronecker-factored-approximate-fisher","repo_url":"https://github.com/Abdoulaye-Koroko/natural-gradients","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"stochastic-optimization","task_name":"Stochastic Optimization"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"sgd","method_name":"SGD"},{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1602.01407","atlas_url":"https://app.syntology.ai/?focus=1602.01407","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1602.01407"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Thrandis/EKFAC-pytorch","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/lzhangbv/eva","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/n-gao/pytorch-kfac","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Abdoulaye-Koroko/natural-gradients","reach":null}],"summary":{"ran_draft_wrong":1},"by_repo_kind":{},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"c95310b1a58f89e9","entry":"grad_wrt_kernel","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"c95310b1a58f89e9"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}