{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/equilibrated-adaptive-learning-rates-for-non","title":"Equilibrated adaptive learning rates for non-convex optimization","arxiv_id":"1502.04390","date":"2015-02-15","proceeding":"NeurIPS 2015 12","authors":["Yann N. Dauphin","Harm de Vries","Yoshua Bengio"],"abstract":"Parameter-specific adaptive learning rate methods are computationally\nefficient ways to reduce the ill-conditioning problems encountered when\ntraining large deep networks. Following recent work that strongly suggests that\nmost of the critical points encountered when training such networks are saddle\npoints, we find how considering the presence of negative eigenvalues of the\nHessian could help us design better suited adaptive learning rate schemes. We\nshow that the popular Jacobi preconditioner has undesirable behavior in the\npresence of both positive and negative curvature, and present theoretical and\nempirical evidence that the so-called equilibration preconditioner is\ncomparatively better suited to non-convex problems. We introduce a novel\nadaptive learning rate scheme, called ESGD, based on the equilibration\npreconditioner. Our experiments show that ESGD performs as well or better than\nRMSProp in terms of convergence speed, always clearly improving over plain\nstochastic gradient descent.","url_abs":"http://arxiv.org/abs/1502.04390v2","url_pdf":"http://arxiv.org/pdf/1502.04390v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"equilibrated-adaptive-learning-rates-for-non","repo_url":"https://github.com/crowsonkb/esgd","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"equilibrated-adaptive-learning-rates-for-non","repo_url":"https://github.com/lixilinx/psgd_tf","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1502.04390","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}