{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/vr-sgd-a-simple-stochastic-variance-reduction","title":"VR-SGD: A Simple Stochastic Variance Reduction Method for Machine Learning","arxiv_id":"1802.09932","date":"2018-02-26","proceeding":null,"authors":["Fanhua Shang","Kaiwen Zhou","Hongying Liu","James Cheng","Ivor W. Tsang","Lijun Zhang","DaCheng Tao","Licheng Jiao"],"abstract":"In this paper, we propose a simple variant of the original SVRG, called\nvariance reduced stochastic gradient descent (VR-SGD). Unlike the choices of\nsnapshot and starting points in SVRG and its proximal variant, Prox-SVRG, the\ntwo vectors of VR-SGD are set to the average and last iterate of the previous\nepoch, respectively. The settings allow us to use much larger learning rates,\nand also make our convergence analysis more challenging. We also design two\ndifferent update rules for smooth and non-smooth objective functions,\nrespectively, which means that VR-SGD can tackle non-smooth and/or non-strongly\nconvex problems directly without any reduction techniques. Moreover, we analyze\nthe convergence properties of VR-SGD for strongly convex problems, which show\nthat VR-SGD attains linear convergence. Different from its counterparts that\nhave no convergence guarantees for non-strongly convex problems, we also\nprovide the convergence guarantees of VR-SGD for this case, and empirically\nverify that VR-SGD with varying learning rates achieves similar performance to\nits momentum accelerated variant that has the optimal convergence rate\n$\\mathcal{O}(1/T^2)$. Finally, we apply VR-SGD to solve various machine\nlearning problems, such as convex and non-convex empirical risk minimization,\nand leading eigenvalue computation. Experimental results show that VR-SGD\nconverges significantly faster than SVRG and Prox-SVRG, and usually outperforms\nstate-of-the-art accelerated methods, e.g., Katyusha.","url_abs":"http://arxiv.org/abs/1802.09932v2","url_pdf":"http://arxiv.org/pdf/1802.09932v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"vr-sgd-a-simple-stochastic-variance-reduction","repo_url":"https://github.com/jnhujnhu/VR-SGD","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"machine-learning","task_name":"BIG-bench Machine Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1802.09932","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}