{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/riemannian-adaptive-optimization-methods","title":"Riemannian Adaptive Optimization Methods","arxiv_id":"1810.00760","date":"2018-10-01","proceeding":"ICLR 2019 5","authors":["Gary Bécigneul","Octavian-Eugen Ganea"],"abstract":"Several first order stochastic optimization methods commonly used in the\nEuclidean domain such as stochastic gradient descent (SGD), accelerated\ngradient descent or variance reduced methods have already been adapted to\ncertain Riemannian settings. However, some of the most popular of these\noptimization tools - namely Adam , Adagrad and the more recent Amsgrad - remain\nto be generalized to Riemannian manifolds. We discuss the difficulty of\ngeneralizing such adaptive schemes to the most agnostic Riemannian setting, and\nthen provide algorithms and convergence proofs for geodesically convex\nobjectives in the particular case of a product of Riemannian manifolds, in\nwhich adaptivity is implemented across manifolds in the cartesian product. Our\ngeneralization is tight in the sense that choosing the Euclidean space as\nRiemannian manifold yields the same algorithms and regret bounds as those that\nwere already known for the standard algorithms. Experimentally, we show faster\nconvergence and to a lower train loss value for Riemannian adaptive methods\nover their corresponding baselines on the realistic task of embedding the\nWordNet taxonomy in the Poincare ball.","url_abs":"http://arxiv.org/abs/1810.00760v2","url_pdf":"http://arxiv.org/pdf/1810.00760v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"riemannian-adaptive-optimization-methods","repo_url":"https://github.com/geoopt/geoopt","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"riemannian-optimization","task_name":"Riemannian optimization"},{"task_slug":"stochastic-optimization","task_name":"Stochastic Optimization"}],"methods":[{"method_slug":"adagrad","method_name":"AdaGrad"},{"method_slug":"adam","method_name":"Adam"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.00760","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}