{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/architectural-complexity-measures-of","title":"Architectural Complexity Measures of Recurrent Neural Networks","arxiv_id":"1602.08210","date":"2016-02-26","proceeding":"NeurIPS 2016 12","authors":["Saizheng Zhang","Yuhuai Wu","Tong Che","Zhouhan Lin","Roland Memisevic","Ruslan Salakhutdinov","Yoshua Bengio"],"abstract":"In this paper, we systematically analyze the connecting architectures of\nrecurrent neural networks (RNNs). Our main contribution is twofold: first, we\npresent a rigorous graph-theoretic framework describing the connecting\narchitectures of RNNs in general. Second, we propose three architecture\ncomplexity measures of RNNs: (a) the recurrent depth, which captures the RNN's\nover-time nonlinear complexity, (b) the feedforward depth, which captures the\nlocal input-output nonlinearity (similar to the \"depth\" in feedforward neural\nnetworks (FNNs)), and (c) the recurrent skip coefficient which captures how\nrapidly the information propagates over time. We rigorously prove each\nmeasure's existence and computability. Our experimental results show that RNNs\nmight benefit from larger recurrent depth and feedforward depth. We further\ndemonstrate that increasing recurrent skip coefficient offers performance\nboosts on long term dependency problems.","url_abs":"http://arxiv.org/abs/1602.08210v3","url_pdf":"http://arxiv.org/pdf/1602.08210v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"language-modelling","task_name":"Language Modelling"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/language-modelling-on-text8","task":"Language Modelling","dataset":"Text8","model":"td-LSTM-large","rank_in_archive_order":23,"of":24,"metrics":{"Bit per Character (BPC)":"1.49"},"uses_additional_data":false},{"leaderboard":"/sota/language-modelling-on-text8","task":"Language Modelling","dataset":"Text8","model":"td-LSTM (Zhang et al., 2016)","rank_in_archive_order":24,"of":24,"metrics":{"Bit per Character (BPC)":"1.63"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1602.08210","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}