{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/scalable-kernel-methods-via-doubly-stochastic","title":"Scalable Kernel Methods via Doubly Stochastic Gradients","arxiv_id":"1407.5599","date":"2014-07-21","proceeding":"NeurIPS 2014 12","authors":["Bo Dai","Bo Xie","Niao He","YIngyu Liang","Anant Raj","Maria-Florina Balcan","Le Song"],"abstract":"The general perception is that kernel methods are not scalable, and neural\nnets are the methods of choice for nonlinear learning problems. Or have we\nsimply not tried hard enough for kernel methods? Here we propose an approach\nthat scales up kernel methods using a novel concept called \"doubly stochastic\nfunctional gradients\". Our approach relies on the fact that many kernel methods\ncan be expressed as convex optimization problems, and we solve the problems by\nmaking two unbiased stochastic approximations to the functional gradient, one\nusing random training points and another using random functions associated with\nthe kernel, and then descending using this noisy functional gradient. We show\nthat a function produced by this procedure after $t$ iterations converges to\nthe optimal function in the reproducing kernel Hilbert space in rate $O(1/t)$,\nand achieves a generalization performance of $O(1/\\sqrt{t})$. This doubly\nstochasticity also allows us to avoid keeping the support vectors and to\nimplement the algorithm in a small memory footprint, which is linear in number\nof iterations and independent of data dimension. Our approach can readily scale\nkernel methods up to the regimes which are dominated by neural nets. We show\nthat our method can achieve competitive performance to neural nets in datasets\nsuch as 8 million handwritten digits from MNIST, 2.3 million energy materials\nfrom MolecularSpace, and 1 million photos from ImageNet.","url_abs":"http://arxiv.org/abs/1407.5599v4","url_pdf":"http://arxiv.org/pdf/1407.5599v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"scalable-kernel-methods-via-doubly-stochastic","repo_url":"https://github.com/zixu1986/Doubly_Stochastic_Gradients","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1407.5599","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}