{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/memory-time-span-in-lstms-for-multi-speaker","title":"Memory Time Span in LSTMs for Multi-Speaker Source Separation","arxiv_id":"1808.08097","date":"2018-08-24","proceeding":null,"authors":["Jeroen Zegers","Hugo Van hamme"],"abstract":"With deep learning approaches becoming state-of-the-art in many speech (as\nwell as non-speech) related machine learning tasks, efforts are being taken to\ndelve into the neural networks which are often considered as a black box. In\nthis paper it is analyzed how recurrent neural network (RNNs) cope with\ntemporal dependencies by determining the relevant memory time span in a long\nshort-term memory (LSTM) cell. This is done by leaking the state variable with\na controlled lifetime and evaluating the task performance. This technique can\nbe used for any task to estimate the time span the LSTM exploits in that\nspecific scenario. The focus in this paper is on the task of separating\nspeakers from overlapping speech. We discern two effects: A long term effect,\nprobably due to speaker characterization and a short term effect, probably\nexploiting phone-size formant tracks.","url_abs":"http://arxiv.org/abs/1808.08097v1","url_pdf":"http://arxiv.org/pdf/1808.08097v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"memory-time-span-in-lstms-for-multi-speaker","repo_url":"https://github.com/JeroenZegers/Nabu-MSSS","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"multi-speaker-source-separation","task_name":"Multi-Speaker Source Separation"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}