{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/diving-into-the-shallows-a-computational","title":"Diving into the shallows: a computational perspective on large-scale shallow learning","arxiv_id":"1703.10622","date":"2017-03-30","proceeding":"NeurIPS 2017 12","authors":["Siyuan Ma","Mikhail Belkin"],"abstract":"In this paper we first identify a basic limitation in gradient descent-based\noptimization methods when used in conjunctions with smooth kernels. An analysis\nbased on the spectral properties of the kernel demonstrates that only a\nvanishingly small portion of the function space is reachable after a polynomial\nnumber of gradient descent iterations. This lack of approximating power\ndrastically limits gradient descent for a fixed computational budget leading to\nserious over-regularization/underfitting. The issue is purely algorithmic,\npersisting even in the limit of infinite data.\n  To address this shortcoming in practice, we introduce EigenPro iteration,\nbased on a preconditioning scheme using a small number of approximately\ncomputed eigenvectors. It can also be viewed as learning a new kernel optimized\nfor gradient descent. It turns out that injecting this small (computationally\ninexpensive and SGD-compatible) amount of approximate second-order information\nleads to major improvements in convergence. For large data, this translates\ninto significant performance boost over the standard kernel methods. In\nparticular, we are able to consistently match or improve the state-of-the-art\nresults recently reported in the literature with a small fraction of their\ncomputational budget.\n  Finally, we feel that these results show a need for a broader computational\nperspective on modern large-scale learning to complement more traditional\nstatistical and convergence analyses. In particular, many phenomena of\nlarge-scale high-dimensional inference are best understood in terms of\noptimization on infinite dimensional Hilbert spaces, where standard algorithms\ncan sometimes have properties at odds with finite-dimensional intuition. A\nsystematic analysis concentrating on the approximation power of such algorithms\nwithin a budget of computation may lead to progress both in theory and\npractice.","url_abs":"http://arxiv.org/abs/1703.10622v2","url_pdf":"http://arxiv.org/pdf/1703.10622v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"diving-into-the-shallows-a-computational","repo_url":"https://github.com/EigenPro/EigenPro-matlab","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1703.10622","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}