{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/meta-learning-with-hessian-free-approach-in","title":"Meta-Learning with Hessian-Free Approach in Deep Neural Nets Training","arxiv_id":"1805.08462","date":"2018-05-22","proceeding":null,"authors":["Boyu Chen","Wenlian Lu","Ernest Fokoue"],"abstract":"Meta-learning is a promising method to achieve efficient training method\ntowards deep neural net and has been attracting increases interests in recent\nyears. But most of the current methods are still not capable to train complex\nneuron net model with long-time training process. In this paper, a novel\nsecond-order meta-optimizer, named Meta-learning with Hessian-Free(MLHF)\napproach, is proposed based on the Hessian-Free approach. Two recurrent neural\nnetworks are established to generate the damping and the precondition matrix of\nthis Hessian-Free framework. A series of techniques to meta-train the MLHF\ntowards stable and reinforce the meta-training of this optimizer, including the\ngradient calculation of $H$. Numerical experiments on deep convolution neural\nnets, including CUDA-convnet and ResNet18(v2), with datasets of CIFAR10 and\nILSVRC2012, indicate that the MLHF shows good and continuous training\nperformance during the whole long-time training process, i.e., both the\nrapid-decreasing early stage and the steadily-deceasing later stage, and so is\na promising meta-learning framework towards elevating the training efficiency\nin real-world deep neural nets.","url_abs":"http://arxiv.org/abs/1805.08462v2","url_pdf":"http://arxiv.org/pdf/1805.08462v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"meta-learning-with-hessian-free-approach-in","repo_url":"https://github.com/ozzzp/MLHF","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"meta-learning","task_name":"Meta-Learning"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}