{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spectral-subspace-dictionary-learning","title":"Dictionary Learning for the Almost-Linear Sparsity Regime","arxiv_id":"2210.10855","date":"2022-10-19","proceeding":null,"authors":["Alexei Novikov","Stephen White"],"abstract":"Dictionary learning, the problem of recovering a sparsely used matrix $\\mathbf{D} \\in \\mathbb{R}^{M \\times K}$ and $N$ $s$-sparse vectors $\\mathbf{x}_i \\in \\mathbb{R}^{K}$ from samples of the form $\\mathbf{y}_i = \\mathbf{D}\\mathbf{x}_i$, is of increasing importance to applications in signal processing and data science. When the dictionary is known, recovery of $\\mathbf{x}_i$ is possible even for sparsity linear in dimension $M$, yet to date, the only algorithms which provably succeed in the linear sparsity regime are Riemannian trust-region methods, which are limited to orthogonal dictionaries, and methods based on the sum-of-squares hierarchy, which requires super-polynomial time in order to obtain an error which decays in $M$. In this work, we introduce SPORADIC (SPectral ORAcle DICtionary Learning), an efficient spectral method on family of reweighted covariance matrices. We prove that in high enough dimensions, SPORADIC can recover overcomplete ($K > M$) dictionaries satisfying the well-known restricted isometry property (RIP) even when sparsity is linear in dimension up to logarithmic factors. Moreover, these accuracy guarantees have an ``oracle property\" that the support and signs of the unknown sparse vectors $\\mathbf{x}_i$ can be recovered exactly with high probability, allowing for arbitrarily close estimation of $\\mathbf{D}$ with enough samples in polynomial time. To the author's knowledge, SPORADIC is the first polynomial-time algorithm which provably enjoys such convergence guarantees for overcomplete RIP matrices in the near-linear sparsity regime.","url_abs":"https://arxiv.org/abs/2210.10855v2","url_pdf":"https://arxiv.org/pdf/2210.10855v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"spectral-subspace-dictionary-learning","repo_url":"https://github.com/sew347/spectral_dict_learn","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"dictionary-learning","task_name":"Dictionary Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}