{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/bayesian-neural-networks-at-finite","title":"Bayesian Neural Networks at Finite Temperature","arxiv_id":"1904.04154","date":"2019-04-08","proceeding":null,"authors":["Robert J. N. Baldock","Nicola Marzari"],"abstract":"We recapitulate the Bayesian formulation of neural network based classifiers\nand show that, while sampling from the posterior does indeed lead to better\ngeneralisation than is obtained by standard optimisation of the cost function,\neven better performance can in general be achieved by sampling finite\ntemperature ($T$) distributions derived from the posterior. Taking the example\nof two different deep (3 hidden layers) classifiers for MNIST data, we find\nquite different $T$ values to be appropriate in each case. In particular, for a\ntypical neural network classifier a clear minimum of the test error is observed\nat $T>0$. This suggests an early stopping criterion for full batch simulated\nannealing: cool until the average validation error starts to increase, then\nrevert to the parameters with the lowest validation error. As $T$ is increased\nclassifiers transition from accurate classifiers to classifiers that have\nhigher training error than assigning equal probability to each class. Efficient\nstudies of these temperature-induced effects are enabled using a\nreplica-exchange Hamiltonian Monte Carlo simulation technique. Finally, we show\nhow thermodynamic integration can be used to perform model selection for deep\nneural networks. Similar to the Laplace approximation, this approach assumes\nthat the posterior is dominated by a single mode. Crucially, however, no\nassumption is made about the shape of that mode and it is not required to\nprecisely compute and invert the Hessian.","url_abs":"http://arxiv.org/abs/1904.04154v1","url_pdf":"http://arxiv.org/pdf/1904.04154v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"bayesian-neural-networks-at-finite","repo_url":"https://github.com/rjnbaldock/nn_sample","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"bayesian-inference","task_name":"Bayesian Inference"},{"task_slug":"model-selection","task_name":"Model Selection"}],"methods":[{"method_slug":"early-stopping","method_name":"Early Stopping"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1904.04154","atlas_url":"https://app.syntology.ai/?focus=1904.04154","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}