{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/approximate-leave-one-out-for-high","title":"Approximate Leave-One-Out for High-Dimensional Non-Differentiable Learning Problems","arxiv_id":"1810.02716","date":"2018-10-04","proceeding":null,"authors":["Shuaiwen Wang","Wenda Zhou","Arian Maleki","Haihao Lu","Vahab Mirrokni"],"abstract":"Consider the following class of learning schemes: \\begin{equation}\n\\label{eq:main-problem1}\n  \\hat{\\boldsymbol{\\beta}} := \\underset{\\boldsymbol{\\beta} \\in\n\\mathcal{C}}{\\arg\\min} \\;\\sum_{j=1}^n\n\\ell(\\boldsymbol{x}_j^\\top\\boldsymbol{\\beta}; y_j) + \\lambda\nR(\\boldsymbol{\\beta}), \\qquad \\qquad \\qquad (1) \\end{equation} where\n$\\boldsymbol{x}_i \\in \\mathbb{R}^p$ and $y_i \\in \\mathbb{R}$ denote the $i^{\\rm\nth}$ feature and response variable respectively. Let $\\ell$ and $R$ be the\nconvex loss function and regularizer, $\\boldsymbol{\\beta}$ denote the unknown\nweights, and $\\lambda$ be a regularization parameter. $\\mathcal{C} \\subset\n\\mathbb{R}^{p}$ is a closed convex set. Finding the optimal choice of $\\lambda$\nis a challenging problem in high-dimensional regimes where both $n$ and $p$ are\nlarge. We propose three frameworks to obtain a computationally efficient\napproximation of the leave-one-out cross validation (LOOCV) risk for nonsmooth\nlosses and regularizers. Our three frameworks are based on the primal, dual,\nand proximal formulations of (1). Each framework shows its strength in certain\ntypes of problems. We prove the equivalence of the three approaches under\nsmoothness conditions. This equivalence enables us to justify the accuracy of\nthe three methods under such conditions. We use our approaches to obtain a risk\nestimate for several standard problems, including generalized LASSO, nuclear\nnorm regularization, and support vector machines. We empirically demonstrate\nthe effectiveness of our results for non-differentiable cases.","url_abs":"http://arxiv.org/abs/1810.02716v1","url_pdf":"http://arxiv.org/pdf/1810.02716v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"approximate-leave-one-out-for-high","repo_url":"https://github.com/wendazhou/alocv-package","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[{"task_slug":"high","task_name":"Vocal Bursts Intensity Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}