{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/learning-to-generalize-provably-in-learning","title":"Learning to Generalize Provably in Learning to Optimize","arxiv_id":"2302.11085","date":"2023-02-22","proceeding":null,"authors":["Junjie Yang","Tianlong Chen","Mingkang Zhu","Fengxiang He","DaCheng Tao","Yingbin Liang","Zhangyang Wang"],"abstract":"Learning to optimize (L2O) has gained increasing popularity, which automates the design of optimizers by data-driven approaches. However, current L2O methods often suffer from poor generalization performance in at least two folds: (i) applying the L2O-learned optimizer to unseen optimizees, in terms of lowering their loss function values (optimizer generalization, or ``generalizable learning of optimizers\"); and (ii) the test performance of an optimizee (itself as a machine learning model), trained by the optimizer, in terms of the accuracy over unseen data (optimizee generalization, or ``learning to generalize\"). While the optimizer generalization has been recently studied, the optimizee generalization (or learning to generalize) has not been rigorously studied in the L2O context, which is the aim of this paper. We first theoretically establish an implicit connection between the local entropy and the Hessian, and hence unify their roles in the handcrafted design of generalizable optimizers as equivalent metrics of the landscape flatness of loss functions. We then propose to incorporate these two metrics as flatness-aware regularizers into the L2O framework in order to meta-train optimizers to learn to generalize, and theoretically show that such generalization ability can be learned during the L2O meta-training process and then transformed to the optimizee loss function. Extensive experiments consistently validate the effectiveness of our proposals with substantially improved generalization on multiple sophisticated L2O models and diverse optimizees. Our code is available at: https://github.com/VITA-Group/Open-L2O/tree/main/Model_Free_L2O/L2O-Entropy.","url_abs":"https://arxiv.org/abs/2302.11085v2","url_pdf":"https://arxiv.org/pdf/2302.11085v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"learning-to-generalize-provably-in-learning","repo_url":"https://github.com/VITA-Group/Open-L2O","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[],"methods":[{"method_slug":"test","method_name":"Test"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2302.11085","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2302.11085"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/VITA-Group/Open-L2O","reach":null}],"summary":{"ran_honours":2},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"5e9f7d2ec2c6eb22","entry":"sample_numiter","repo":"VITA-Group/Open-L2O","repo_kind":"official","path":"Model_Free_L2O/L2O-Entropy/L2O-ScalewHessian/L2O-Scale-Evaluation/metaopt.py","file_url":"https://github.com/VITA-Group/Open-L2O/blob/HEAD/Model_Free_L2O/L2O-Entropy/L2O-ScalewHessian/L2O-Scale-Evaluation/metaopt.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"5e9f7d2ec2c6eb22"}},{"code_sha256_prefix":"4bf86510c847be7c","entry":"sigmoid_weights","repo":"VITA-Group/Open-L2O","repo_kind":"official","path":"Model_Free_L2O/L2O-Entropy/L2O-ScalewHessian/L2O-Scale-Evaluation/metaopt.py","file_url":"https://github.com/VITA-Group/Open-L2O/blob/HEAD/Model_Free_L2O/L2O-Entropy/L2O-ScalewHessian/L2O-Scale-Evaluation/metaopt.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"4bf86510c847be7c"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}