{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spectrum-estimation-from-samples","title":"Spectrum Estimation from Samples","arxiv_id":"1602.00061","date":"2016-01-30","proceeding":null,"authors":["Weihao Kong","Gregory Valiant"],"abstract":"We consider the problem of approximating the set of eigenvalues of the\ncovariance matrix of a multivariate distribution (equivalently, the problem of\napproximating the \"population spectrum\"), given access to samples drawn from\nthe distribution. The eigenvalues of the covariance of a distribution contain\nbasic information about the distribution, including the presence or lack of\nstructure in the distribution, the effective dimensionality of the\ndistribution, and the applicability of higher-level machine learning and\nmultivariate statistical tools. We consider this fundamental recovery problem\nin the regime where the number of samples is comparable, or even sublinear in\nthe dimensionality of the distribution in question. First, we propose a\ntheoretically optimal and computationally efficient algorithm for recovering\nthe moments of the eigenvalues of the population covariance matrix. We then\nleverage this accurate moment recovery, via a Wasserstein distance argument, to\nshow that the vector of eigenvalues can be accurately recovered. We provide\nfinite--sample bounds on the expected error of the recovered eigenvalues, which\nimply that our estimator is asymptotically consistent as the dimensionality of\nthe distribution and sample size tend towards infinity, even in the sublinear\nsample regime where the ratio of the sample size to the dimensionality tends to\nzero. In addition to our theoretical results, we show that our approach\nperforms well in practice for a broad range of distributions and sample sizes.","url_abs":"http://arxiv.org/abs/1602.00061v5","url_pdf":"http://arxiv.org/pdf/1602.00061v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"spectrum-estimation-from-samples","repo_url":"https://github.com/harinath001/compbio-project","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1602.00061","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}