{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-robust-adaptive-stochastic-gradient-method","title":"A Robust Adaptive Stochastic Gradient Method for Deep Learning","arxiv_id":"1703.00788","date":"2017-03-02","proceeding":null,"authors":["Caglar Gulcehre","Jose Sotelo","Marcin Moczulski","Yoshua Bengio"],"abstract":"Stochastic gradient algorithms are the main focus of large-scale optimization\nproblems and led to important successes in the recent advancement of the deep\nlearning algorithms. The convergence of SGD depends on the careful choice of\nlearning rate and the amount of the noise in stochastic estimates of the\ngradients. In this paper, we propose an adaptive learning rate algorithm, which\nutilizes stochastic curvature information of the loss function for\nautomatically tuning the learning rates. The information about the element-wise\ncurvature of the loss function is estimated from the local statistics of the\nstochastic first order gradients. We further propose a new variance reduction\ntechnique to speed up the convergence. In our experiments with deep neural\nnetworks, we obtained better performance compared to the popular stochastic\ngradient algorithms.","url_abs":"http://arxiv.org/abs/1703.00788v1","url_pdf":"http://arxiv.org/pdf/1703.00788v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-robust-adaptive-stochastic-gradient-method","repo_url":"https://github.com/sotelo/scribe","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"deep-learning","task_name":"Deep Learning"}],"methods":[{"method_slug":"sgd","method_name":"SGD"},{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}