{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-the-long-term-memory-of-deep-recurrent","title":"On the Long-Term Memory of Deep Recurrent Networks","arxiv_id":"1710.09431","date":"2017-10-25","proceeding":null,"authors":["Yoav Levine","Or Sharir","Alon Ziv","Amnon Shashua"],"abstract":"A key attribute that drives the unprecedented success of modern Recurrent\nNeural Networks (RNNs) on learning tasks which involve sequential data, is\ntheir ability to model intricate long-term temporal dependencies. However, a\nwell established measure of RNNs long-term memory capacity is lacking, and thus\nformal understanding of the effect of depth on their ability to correlate data\nthroughout time is limited. Specifically, existing depth efficiency results on\nconvolutional networks do not suffice in order to account for the success of\ndeep RNNs on data of varying lengths. In order to address this, we introduce a\nmeasure of the network's ability to support information flow across time,\nreferred to as the Start-End separation rank, which reflects the distance of\nthe function realized by the recurrent network from modeling no dependency\nbetween the beginning and end of the input sequence. We prove that deep\nrecurrent networks support Start-End separation ranks which are combinatorially\nhigher than those supported by their shallow counterparts. Thus, we establish\nthat depth brings forth an overwhelming advantage in the ability of recurrent\nnetworks to model long-term dependencies, and provide an exemplar of\nquantifying this key attribute which may be readily extended to other RNN\narchitectures of interest, e.g. variants of LSTM networks. We obtain our\nresults by considering a class of recurrent networks referred to as Recurrent\nArithmetic Circuits, which merge the hidden state with the input via the\nMultiplicative Integration operation, and empirically demonstrate the discussed\nphenomena on common RNNs. Finally, we employ the tool of quantum Tensor\nNetworks to gain additional graphic insight regarding the complexity brought\nforth by depth in recurrent networks.","url_abs":"http://arxiv.org/abs/1710.09431v2","url_pdf":"http://arxiv.org/pdf/1710.09431v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-the-long-term-memory-of-deep-recurrent","repo_url":"https://github.com/HUJI-Deep/Long-Term-Memory-of-Deep-RNNs","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"tensor-networks","task_name":"Tensor Networks"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1710.09431","atlas_url":"https://app.syntology.ai/?focus=1710.09431","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}