{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/orthogonal-weight-normalization-solution-to","title":"Orthogonal Weight Normalization: Solution to Optimization over Multiple Dependent Stiefel Manifolds in Deep Neural Networks","arxiv_id":"1709.06079","date":"2017-09-16","proceeding":null,"authors":["Lei Huang","Xianglong Liu","Bo Lang","Adams Wei Yu","Yongliang Wang","Bo Li"],"abstract":"Orthogonal matrix has shown advantages in training Recurrent Neural Networks\n(RNNs), but such matrix is limited to be square for the hidden-to-hidden\ntransformation in RNNs. In this paper, we generalize such square orthogonal\nmatrix to orthogonal rectangular matrix and formulating this problem in\nfeed-forward Neural Networks (FNNs) as Optimization over Multiple Dependent\nStiefel Manifolds (OMDSM). We show that the rectangular orthogonal matrix can\nstabilize the distribution of network activations and regularize FNNs. We also\npropose a novel orthogonal weight normalization method to solve OMDSM.\nParticularly, it constructs orthogonal transformation over proxy parameters to\nensure the weight matrix is orthogonal and back-propagates gradient information\nthrough the transformation during training. To guarantee stability, we minimize\nthe distortions between proxy parameters and canonical weights over all\ntractable orthogonal transformations. In addition, we design an orthogonal\nlinear module (OLM) to learn orthogonal filter banks in practice, which can be\nused as an alternative to standard linear module. Extensive experiments\ndemonstrate that by simply substituting OLM for standard linear module without\nrevising any experimental protocols, our method largely improves the\nperformance of the state-of-the-art networks, including Inception and residual\nnetworks on CIFAR and ImageNet datasets. In particular, we have reduced the\ntest error of wide residual network on CIFAR-100 from 20.04% to 18.61% with\nsuch simple substitution. Our code is available online for result reproduction.","url_abs":"http://arxiv.org/abs/1709.06079v2","url_pdf":"http://arxiv.org/pdf/1709.06079v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"orthogonal-weight-normalization-solution-to","repo_url":"https://github.com/huangleiBuaa/OthogonalWN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"image-classification","task_name":"Image Classification"}],"methods":[{"method_slug":"weight-normalization","method_name":"Weight Normalization"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1709.06079","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}