{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-orthogonality-and-learning-recurrent","title":"On orthogonality and learning recurrent networks with long term dependencies","arxiv_id":"1702.00071","date":"2017-01-31","proceeding":null,"authors":["Eugene Vorontsov","Chiheb Trabelsi","Samuel Kadoury","Chris Pal"],"abstract":"It is well known that it is challenging to train deep neural networks and\nrecurrent neural networks for tasks that exhibit long term dependencies. The\nvanishing or exploding gradient problem is a well known issue associated with\nthese challenges. One approach to addressing vanishing and exploding gradients\nis to use either soft or hard constraints on weight matrices so as to encourage\nor enforce orthogonality. Orthogonal matrices preserve gradient norm during\nbackpropagation and may therefore be a desirable property. This paper explores\nissues with optimization convergence, speed and gradient stability when\nencouraging or enforcing orthogonality. To perform this analysis, we propose a\nweight matrix factorization and parameterization strategy through which we can\nbound matrix norms and therein control the degree of expansivity induced during\nbackpropagation. We find that hard constraints on orthogonality can negatively\naffect the speed of convergence and model performance.","url_abs":"http://arxiv.org/abs/1702.00071v4","url_pdf":"http://arxiv.org/pdf/1702.00071v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-orthogonality-and-learning-recurrent","repo_url":"https://github.com/veugene/spectre_release","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[],"methods":[{"method_slug":"speed","method_name":"SPEED"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1702.00071","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}