{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/kernel-machines-that-adapt-to-gpus-for","title":"Kernel machines that adapt to GPUs for effective large batch training","arxiv_id":"1806.06144","date":"2018-06-15","proceeding":null,"authors":["Siyuan Ma","Mikhail Belkin"],"abstract":"Modern machine learning models are typically trained using Stochastic\nGradient Descent (SGD) on massively parallel computing resources such as GPUs.\nIncreasing mini-batch size is a simple and direct way to utilize the parallel\ncomputing capacity. For small batch an increase in batch size results in the\nproportional reduction in the training time, a phenomenon known as linear\nscaling. However, increasing batch size beyond a certain value leads to no\nfurther improvement in training time. In this paper we develop the first\nanalytical framework that extends linear scaling to match the parallel\ncomputing capacity of a resource. The framework is designed for a class of\nclassical kernel machines. It automatically modifies a standard kernel machine\nto output a mathematically equivalent prediction function, yet allowing for\nextended linear scaling, i.e., higher effective parallelization and faster\ntraining time on given hardware.\n  The resulting algorithms are accurate, principled and very fast. For example,\nusing a single Titan Xp GPU, training on ImageNet with $1.3\\times 10^6$ data\npoints and $1000$ labels takes under an hour, while smaller datasets, such as\nMNIST, take seconds. As the parameters are chosen analytically, based on the\ntheoretical bounds, little tuning beyond selecting the kernel and the kernel\nparameter is needed, further facilitating the practical use of these methods.","url_abs":"http://arxiv.org/abs/1806.06144v3","url_pdf":"http://arxiv.org/pdf/1806.06144v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"kernel-machines-that-adapt-to-gpus-for","repo_url":"https://github.com/EigenPro/EigenPro2","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"kernel-machines-that-adapt-to-gpus-for","repo_url":"https://github.com/eigenpro/eigenpro-pytorch","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":null,"task_name":"GPU"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}