{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/normalized-direction-preserving-adam","title":"Normalized Direction-preserving Adam","arxiv_id":"1709.04546","date":"2017-09-13","proceeding":"ICLR 2018 1","authors":["Zijun Zhang","Lin Ma","Zongpeng Li","Chuan Wu"],"abstract":"Adaptive optimization algorithms, such as Adam and RMSprop, have shown better\noptimization performance than stochastic gradient descent (SGD) in some\nscenarios. However, recent studies show that they often lead to worse\ngeneralization performance than SGD, especially for training deep neural\nnetworks (DNNs). In this work, we identify the reasons that Adam generalizes\nworse than SGD, and develop a variant of Adam to eliminate the generalization\ngap. The proposed method, normalized direction-preserving Adam (ND-Adam),\nenables more precise control of the direction and step size for updating weight\nvectors, leading to significantly improved generalization performance.\nFollowing a similar rationale, we further improve the generalization\nperformance in classification tasks by regularizing the softmax logits. By\nbridging the gap between SGD and Adam, we also hope to shed light on why\ncertain optimization algorithms generalize better than others.","url_abs":"http://arxiv.org/abs/1709.04546v2","url_pdf":"http://arxiv.org/pdf/1709.04546v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"normalized-direction-preserving-adam","repo_url":"https://github.com/zj10/ND-Adam","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"}],"methods":[{"method_slug":"adam","method_name":"Adam"},{"method_slug":"sgd","method_name":"SGD"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1709.04546","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}