{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/outlier-robust-sparse-low-rank-least-squares","title":"Outlier-robust sparse/low-rank least-squares regression and robust matrix completion","arxiv_id":"2012.06750","date":"2020-12-12","proceeding":null,"authors":["Philip Thompson"],"abstract":"We study high-dimensional least-squares regression within a subgaussian statistical learning framework with heterogeneous noise. It includes $s$-sparse and $r$-low-rank least-squares regression when a fraction $\\epsilon$ of the labels are adversarially contaminated. We also present a novel theory of trace-regression with matrix decomposition based on a new application of the product process. For these problems, we show novel near-optimal \"subgaussian\" estimation rates of the form $r(n,d_{e})+\\sqrt{\\log(1/\\delta)/n}+\\epsilon\\log(1/\\epsilon)$, valid with probability at least $1-\\delta$. Here, $r(n,d_{e})$ is the optimal uncontaminated rate as a function of the effective dimension $d_{e}$ but independent of the failure probability $\\delta$. These rates are valid uniformly on $\\delta$, i.e., the estimators' tuning do not depend on $\\delta$. Lastly, we consider noisy robust matrix completion with non-uniform sampling. If only the low-rank matrix is of interest, we present a novel near-optimal rate that is independent of the corruption level $a$. Our estimators are tractable and based on a new \"sorted\" Huber-type loss. No information on $(s,r,\\epsilon,a)$ are needed to tune these estimators. Our analysis makes use of novel $\\delta$-optimal concentration inequalities for the multiplier and product processes which could be useful elsewhere. For instance, they imply novel sharp oracle inequalities for Lasso and Slope with optimal dependence on $\\delta$. Numerical simulations confirm our theoretical predictions. In particular, \"sorted\" Huber regression can outperform classical Huber regression.","url_abs":"https://arxiv.org/abs/2012.06750v3","url_pdf":"https://arxiv.org/pdf/2012.06750v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"outlier-robust-sparse-low-rank-least-squares","repo_url":"https://github.com/philipthomp/Outlier-robust-regression","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"matrix-completion","task_name":"Matrix Completion"},{"task_slug":"regression-1","task_name":"regression"},{"task_slug":null,"task_name":"valid"}],"methods":[{"method_slug":"linear-regression","method_name":"Linear Regression"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}