{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/implicit-self-regularization-in-deep-neural","title":"Implicit Self-Regularization in Deep Neural Networks: Evidence from Random Matrix Theory and Implications for Learning","arxiv_id":"1810.01075","date":"2018-10-02","proceeding":null,"authors":["Charles H. Martin","Michael W. Mahoney"],"abstract":"Random Matrix Theory (RMT) is applied to analyze weight matrices of Deep\nNeural Networks (DNNs), including both production quality, pre-trained models\nsuch as AlexNet and Inception, and smaller models trained from scratch, such as\nLeNet5 and a miniature-AlexNet. Empirical and theoretical results clearly\nindicate that the DNN training process itself implicitly implements a form of\nSelf-Regularization. The empirical spectral density (ESD) of DNN layer matrices\ndisplays signatures of traditionally-regularized statistical models, even in\nthe absence of exogenously specifying traditional forms of explicit\nregularization. Building on relatively recent results in RMT, most notably its\nextension to Universality classes of Heavy-Tailed matrices, we develop a theory\nto identify 5+1 Phases of Training, corresponding to increasing amounts of\nImplicit Self-Regularization. These phases can be observed during the training\nprocess as well as in the final learned DNNs. For smaller and/or older DNNs,\nthis Implicit Self-Regularization is like traditional Tikhonov regularization,\nin that there is a \"size scale\" separating signal from noise. For\nstate-of-the-art DNNs, however, we identify a novel form of Heavy-Tailed\nSelf-Regularization, similar to the self-organization seen in the statistical\nphysics of disordered systems. This results from correlations arising at all\nsize scales, which arises implicitly due to the training process itself. This\nimplicit Self-Regularization can depend strongly on the many knobs of the\ntraining process. By exploiting the generalization gap phenomena, we\ndemonstrate that we can cause a small model to exhibit all 5+1 phases of\ntraining simply by changing the batch size. This demonstrates that---all else\nbeing equal---DNN optimization with larger batch sizes leads to less-well\nimplicitly-regularized models, and it provides an explanation for the\ngeneralization gap phenomena.","url_abs":"http://arxiv.org/abs/1810.01075v1","url_pdf":"http://arxiv.org/pdf/1810.01075v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"implicit-self-regularization-in-deep-neural","repo_url":"https://github.com/CalculatedContent/ImplicitSelfRegularization","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}},{"paper_slug":"implicit-self-regularization-in-deep-neural","repo_url":"https://github.com/Zachary-Stone-Berkeley/Learning_Rate_Experiments","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"implicit-self-regularization-in-deep-neural","repo_url":"https://github.com/Zachary-Stone/Learning_Rate_Experiments","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"grouped-convolution","method_name":"Grouped Convolution"},{"method_slug":"local-response-normalization","method_name":"Local Response Normalization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.01075","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}