{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/entropy-sgd-biasing-gradient-descent-into","title":"Entropy-SGD: Biasing Gradient Descent Into Wide Valleys","arxiv_id":"1611.01838","date":"2016-11-06","proceeding":null,"authors":["Pratik Chaudhari","Anna Choromanska","Stefano Soatto","Yann Lecun","Carlo Baldassi","Christian Borgs","Jennifer Chayes","Levent Sagun","Riccardo Zecchina"],"abstract":"This paper proposes a new optimization algorithm called Entropy-SGD for\ntraining deep neural networks that is motivated by the local geometry of the\nenergy landscape. Local extrema with low generalization error have a large\nproportion of almost-zero eigenvalues in the Hessian with very few positive or\nnegative eigenvalues. We leverage upon this observation to construct a\nlocal-entropy-based objective function that favors well-generalizable solutions\nlying in large flat regions of the energy landscape, while avoiding\npoorly-generalizable solutions located in the sharp valleys. Conceptually, our\nalgorithm resembles two nested loops of SGD where we use Langevin dynamics in\nthe inner loop to compute the gradient of the local entropy before each update\nof the weights. We show that the new objective has a smoother energy landscape\nand show improved generalization over SGD using uniform stability, under\ncertain assumptions. Our experiments on convolutional and recurrent networks\ndemonstrate that Entropy-SGD compares favorably to state-of-the-art techniques\nin terms of generalization error and training time.","url_abs":"http://arxiv.org/abs/1611.01838v5","url_pdf":"http://arxiv.org/pdf/1611.01838v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"entropy-sgd-biasing-gradient-descent-into","repo_url":"https://github.com/ucla-vision/entropy-sgd","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"entropy-sgd-biasing-gradient-descent-into","repo_url":"https://github.com/steph1793/Entropy-SGD","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[],"methods":[{"method_slug":"sgd","method_name":"SGD"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1611.01838","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}