{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/m-lo-compute-efficient-meta-generalization-of","title":"$μ$LO: Compute-Efficient Meta-Generalization of Learned Optimizers","arxiv_id":"2406.00153","date":"2024-05-31","proceeding":null,"authors":["Benjamin Thérien","Charles-Étienne Joseph","Boris Knyazev","Edouard Oyallon","Irina Rish","Eugene Belilovsky"],"abstract":"Learned optimizers (LOs) can significantly reduce the wall-clock training time of neural networks, substantially reducing training costs. However, they often suffer from poor meta-generalization, especially when training networks larger than those seen during meta-training. To address this, we use the recently proposed Maximal Update Parametrization ($\\mu$P), which allows zero-shot generalization of optimizer hyperparameters from smaller to larger models. We extend $\\mu$P theory to learned optimizers, treating the meta-training problem as finding the learned optimizer under $\\mu$P. Our evaluation shows that LOs meta-trained with $\\mu$P substantially improve meta-generalization as compared to LOs trained under standard parametrization (SP). Notably, when applied to large-width models, our best $\\mu$LO, trained for 103 GPU-hours, matches or exceeds the performance of VeLO, the largest publicly available learned optimizer, meta-trained with 4000 TPU-months of compute. Moreover, $\\mu$LOs demonstrate better generalization than their SP counterparts to deeper networks and to much longer training horizons (25 times longer) than those seen during meta-training.","url_abs":"https://arxiv.org/abs/2406.00153v1","url_pdf":"https://arxiv.org/pdf/2406.00153v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"m-lo-compute-efficient-meta-generalization-of","repo_url":"https://github.com/bentherien/mu_learned_optimization","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"jax","reach":{"status":"ok"}}],"tasks":[{"task_slug":null,"task_name":"GPU"},{"task_slug":"zero-shot-generalization","task_name":"Zero-shot Generalization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.00153","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}