{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/optimization-methods-for-supervised-machine","title":"Optimization Methods for Supervised Machine Learning: From Linear Models to Deep Learning","arxiv_id":"1706.10207","date":"2017-06-30","proceeding":null,"authors":["Frank E. Curtis","Katya Scheinberg"],"abstract":"The goal of this tutorial is to introduce key models, algorithms, and open\nquestions related to the use of optimization methods for solving problems\narising in machine learning. It is written with an INFORMS audience in mind,\nspecifically those readers who are familiar with the basics of optimization\nalgorithms, but less familiar with machine learning. We begin by deriving a\nformulation of a supervised learning problem and show how it leads to various\noptimization problems, depending on the context and underlying assumptions. We\nthen discuss some of the distinctive features of these optimization problems,\nfocusing on the examples of logistic regression and the training of deep neural\nnetworks. The latter half of the tutorial focuses on optimization algorithms,\nfirst for convex logistic regression, for which we discuss the use of\nfirst-order methods, the stochastic gradient method, variance reducing\nstochastic methods, and second-order methods. Finally, we discuss how these\napproaches can be employed to the training of deep neural networks, emphasizing\nthe difficulties that arise from the complex, nonconvex structure of these\nmodels.","url_abs":"http://arxiv.org/abs/1706.10207v1","url_pdf":"http://arxiv.org/pdf/1706.10207v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"optimization-methods-for-supervised-machine","repo_url":"https://github.com/GCaptainNemo/optimization-project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"},{"task_slug":"second-order-methods","task_name":"Second-order methods"},{"task_slug":"regression-1","task_name":"regression"}],"methods":[{"method_slug":"logistic-regression","method_name":"Logistic Regression"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}