{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/when-will-gradient-methods-converge-to-max","title":"When Will Gradient Methods Converge to Max-margin Classifier under ReLU Models?","arxiv_id":"1806.04339","date":"2018-06-12","proceeding":"ICLR 2019 5","authors":["Tengyu Xu","Yi Zhou","Kaiyi Ji","Yingbin Liang"],"abstract":"We study the implicit bias of gradient descent methods in solving a binary\nclassification problem over a linearly separable dataset. The classifier is\ndescribed by a nonlinear ReLU model and the objective function adopts the\nexponential loss function. We first characterize the landscape of the loss\nfunction and show that there can exist spurious asymptotic local minima besides\nasymptotic global minima. We then show that gradient descent (GD) can converge\nto either a global or a local max-margin direction, or may diverge from the\ndesired max-margin direction in a general context. For stochastic gradient\ndescent (SGD), we show that it converges in expectation to either the global or\nthe local max-margin direction if SGD converges. We further explore the\nimplicit bias of these algorithms in learning a multi-neuron network under\ncertain stationary conditions, and show that the learned classifier maximizes\nthe margins of each sample pattern partition under the ReLU activation.","url_abs":"http://arxiv.org/abs/1806.04339v2","url_pdf":"http://arxiv.org/pdf/1806.04339v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"when-will-gradient-methods-converge-to-max","repo_url":"https://github.com/WilliamLiPro/LpSS","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"binary-classification","task_name":"Binary Classification"}],"methods":[{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"sgd","method_name":"SGD"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1806.04339","atlas_url":"https://app.syntology.ai/?focus=1806.04339","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}