{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-quenching-activation-behavior-of-the","title":"The Quenching-Activation Behavior of the Gradient Descent Dynamics for Two-layer Neural Network Models","arxiv_id":"2006.14450","date":"2020-06-25","proceeding":null,"authors":["Chao Ma","Lei Wu","Weinan E"],"abstract":"A numerical and phenomenological study of the gradient descent (GD) algorithm for training two-layer neural network models is carried out for different parameter regimes when the target function can be accurately approximated by a relatively small number of neurons. It is found that for Xavier-like initialization, there are two distinctive phases in the dynamic behavior of GD in the under-parametrized regime: An early phase in which the GD dynamics follows closely that of the corresponding random feature model and the neurons are effectively quenched, followed by a late phase in which the neurons are divided into two groups: a group of a few \"activated\" neurons that dominate the dynamics and a group of background (or \"quenched\") neurons that support the continued activation and deactivation process. This neural network-like behavior is continued into the mildly over-parametrized regime, where it undergoes a transition to a random feature-like behavior. The quenching-activation process seems to provide a clear mechanism for \"implicit regularization\". This is qualitatively different from the dynamics associated with the \"mean-field\" scaling where all neurons participate equally and there does not appear to be qualitative changes when the network parameters are changed.","url_abs":"https://arxiv.org/abs/2006.14450v1","url_pdf":"https://arxiv.org/pdf/2006.14450v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-quenching-activation-behavior-of-the","repo_url":"https://github.com/TheoreticalML/GD.quenching_activation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2006.14450","atlas_url":"https://app.syntology.ai/?focus=2006.14450","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}