{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-learn-stochastic-gradient-descent","title":"Learning-to-Learn Stochastic Gradient Descent with Biased Regularization","arxiv_id":"1903.10399","date":"2019-03-25","proceeding":null,"authors":["Giulia Denevi","Carlo Ciliberto","Riccardo Grazzi","Massimiliano Pontil"],"abstract":"We study the problem of learning-to-learn: inferring a learning algorithm\nthat works well on tasks sampled from an unknown distribution. As class of\nalgorithms we consider Stochastic Gradient Descent on the true risk regularized\nby the square euclidean distance to a bias vector. We present an average excess\nrisk bound for such a learning algorithm. This result quantifies the potential\nbenefit of using a bias vector with respect to the unbiased case. We then\naddress the problem of estimating the bias from a sequence of tasks. We propose\na meta-algorithm which incrementally updates the bias, as new tasks are\nobserved. The low space and time complexity of this approach makes it appealing\nin practice. We provide guarantees on the learning ability of the\nmeta-algorithm. A key feature of our results is that, when the number of tasks\ngrows and their variance is relatively small, our learning-to-learn approach\nhas a significant advantage over learning each task in isolation by Stochastic\nGradient Descent without a bias term. We report on numerical experiments which\ndemonstrate the effectiveness of our approach.","url_abs":"http://arxiv.org/abs/1903.10399v1","url_pdf":"http://arxiv.org/pdf/1903.10399v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-learn-stochastic-gradient-descent","repo_url":"https://github.com/prolearner/onlineLTL","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1903.10399","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}