{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/training-deep-networks-without-learning-rates","title":"Training Deep Networks without Learning Rates Through Coin Betting","arxiv_id":"1705.07795","date":"2017-05-22","proceeding":"NeurIPS 2017 12","authors":["Francesco Orabona","Tatiana Tommasi"],"abstract":"Deep learning methods achieve state-of-the-art performance in many\napplication scenarios. Yet, these methods require a significant amount of\nhyperparameters tuning in order to achieve the best results. In particular,\ntuning the learning rates in the stochastic optimization process is still one\nof the main bottlenecks. In this paper, we propose a new stochastic gradient\ndescent procedure for deep networks that does not require any learning rate\nsetting. Contrary to previous methods, we do not adapt the learning rates nor\nwe make use of the assumed curvature of the objective function. Instead, we\nreduce the optimization process to a game of betting on a coin and propose a\nlearning-rate-free optimal algorithm for this scenario. Theoretical convergence\nis proven for convex and quasi-convex functions and empirical evidence shows\nthe advantage of our algorithm over popular stochastic gradient algorithms.","url_abs":"http://arxiv.org/abs/1705.07795v3","url_pdf":"http://arxiv.org/pdf/1705.07795v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"training-deep-networks-without-learning-rates","repo_url":"https://github.com/bremen79/cocob","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"training-deep-networks-without-learning-rates","repo_url":"https://github.com/anandsaha/nips.cocob.pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"training-deep-networks-without-learning-rates","repo_url":"https://github.com/bremen79/parameterfree","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"training-deep-networks-without-learning-rates","repo_url":"https://github.com/brendanxwhitaker/spred","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"training-deep-networks-without-learning-rates","repo_url":"https://github.com/nocotan/cocob_backprop","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"training-deep-networks-without-learning-rates","repo_url":"https://github.com/tensorflow/addons/blob/master/tensorflow_addons/optimizers/cocob.py","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"stochastic-optimization","task_name":"Stochastic Optimization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/stochastic-optimization-on-mnist","task":"Stochastic Optimization","dataset":"MNIST","model":"MLP","rank_in_archive_order":1,"of":1,"metrics":{"NLL":"0.0541"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}