{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/optimization-methods-for-large-scale-machine","title":"Optimization Methods for Large-Scale Machine Learning","arxiv_id":"1606.04838","date":"2016-06-15","proceeding":null,"authors":["Léon Bottou","Frank E. Curtis","Jorge Nocedal"],"abstract":"This paper provides a review and commentary on the past, present, and future\nof numerical optimization algorithms in the context of machine learning\napplications. Through case studies on text classification and the training of\ndeep neural networks, we discuss how optimization problems arise in machine\nlearning and what makes them challenging. A major theme of our study is that\nlarge-scale machine learning represents a distinctive setting in which the\nstochastic gradient (SG) method has traditionally played a central role while\nconventional gradient-based nonlinear optimization techniques typically falter.\nBased on this viewpoint, we present a comprehensive theory of a\nstraightforward, yet versatile SG algorithm, discuss its practical behavior,\nand highlight opportunities for designing algorithms with improved performance.\nThis leads to a discussion about the next generation of optimization methods\nfor large-scale machine learning, including an investigation of two main\nstreams of research on techniques that diminish noise in the stochastic\ndirections and methods that make use of second-order derivative approximations.","url_abs":"http://arxiv.org/abs/1606.04838v3","url_pdf":"http://arxiv.org/pdf/1606.04838v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"optimization-methods-for-large-scale-machine","repo_url":"https://github.com/Coolgiserz/NLP_starter","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"optimization-methods-for-large-scale-machine","repo_url":"https://github.com/GCaptainNemo/optimization-project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"optimization-methods-for-large-scale-machine","repo_url":"https://github.com/stephenbeckr/AIMS","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"optimization-methods-for-large-scale-machine","repo_url":"https://github.com/stephenbeckr/CambridgeOptimisationCourse","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"text-classification","task_name":"Text Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1606.04838","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}