{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/empirical-risk-minimization-for-stochastic","title":"Empirical Risk Minimization for Stochastic Convex Optimization: $O(1/n)$- and $O(1/n^2)$-type of Risk Bounds","arxiv_id":"1702.02030","date":"2017-02-07","proceeding":null,"authors":["Lijun Zhang","Tianbao Yang","Rong Jin"],"abstract":"Although there exist plentiful theories of empirical risk minimization (ERM)\nfor supervised learning, current theoretical understandings of ERM for a\nrelated problem---stochastic convex optimization (SCO), are limited. In this\nwork, we strengthen the realm of ERM for SCO by exploiting smoothness and\nstrong convexity conditions to improve the risk bounds. First, we establish an\n$\\widetilde{O}(d/n + \\sqrt{F_*/n})$ risk bound when the random function is\nnonnegative, convex and smooth, and the expected function is Lipschitz\ncontinuous, where $d$ is the dimensionality of the problem, $n$ is the number\nof samples, and $F_*$ is the minimal risk. Thus, when $F_*$ is small we obtain\nan $\\widetilde{O}(d/n)$ risk bound, which is analogous to the\n$\\widetilde{O}(1/n)$ optimistic rate of ERM for supervised learning. Second, if\nthe objective function is also $\\lambda$-strongly convex, we prove an\n$\\widetilde{O}(d/n + \\kappa F_*/n )$ risk bound where $\\kappa$ is the condition\nnumber, and improve it to $O(1/[\\lambda n^2] + \\kappa F_*/n)$ when\n$n=\\widetilde{\\Omega}(\\kappa d)$. As a result, we obtain an $O(\\kappa/n^2)$\nrisk bound under the condition that $n$ is large and $F_*$ is small, which to\nthe best of our knowledge, is the first $O(1/n^2)$-type of risk bound of ERM.\nThird, we stress that the above results are established in a unified framework,\nwhich allows us to derive new risk bounds under weaker conditions, e.g.,\nwithout convexity of the random function and Lipschitz continuity of the\nexpected function. Finally, we demonstrate that to achieve an $O(1/[\\lambda\nn^2] + \\kappa F_*/n)$ risk bound for supervised learning, the\n$\\widetilde{\\Omega}(\\kappa d)$ requirement on $n$ can be replaced with\n$\\Omega(\\kappa^2)$, which is dimensionality-independent.","url_abs":"http://arxiv.org/abs/1702.02030v1","url_pdf":"http://arxiv.org/pdf/1702.02030v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/image-classification-on-colored-mnist-with","task":"Image Classification","dataset":"Colored-MNIST(with spurious correlation)","model":"MLP-ERM","rank_in_archive_order":5,"of":6,"metrics":{"Accuracy ":"17.10"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1702.02030","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}