{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-recurrent-nmf-for-speech-separation-by","title":"Deep Recurrent NMF for Speech Separation by Unfolding Iterative Thresholding","arxiv_id":"1709.07124","date":"2017-09-21","proceeding":null,"authors":["Scott Wisdom","Thomas Powers","James Pitton","Les Atlas"],"abstract":"In this paper, we propose a novel recurrent neural network architecture for\nspeech separation. This architecture is constructed by unfolding the iterations\nof a sequential iterative soft-thresholding algorithm (ISTA) that solves the\noptimization problem for sparse nonnegative matrix factorization (NMF) of\nspectrograms. We name this network architecture deep recurrent NMF (DR-NMF).\nThe proposed DR-NMF network has three distinct advantages. First, DR-NMF\nprovides better interpretability than other deep architectures, since the\nweights correspond to NMF model parameters, even after training. This\ninterpretability also provides principled initializations that enable faster\ntraining and convergence to better solutions compared to conventional random\ninitialization. Second, like many deep networks, DR-NMF is an order of\nmagnitude faster at test time than NMF, since computation of the network output\nonly requires evaluating a few layers at each time step. Third, when a limited\namount of training data is available, DR-NMF exhibits stronger generalization\nand separation performance compared to sparse NMF and state-of-the-art\nlong-short term memory (LSTM) networks. When a large amount of training data is\navailable, DR-NMF achieves lower yet competitive separation performance\ncompared to LSTM networks.","url_abs":"http://arxiv.org/abs/1709.07124v1","url_pdf":"http://arxiv.org/pdf/1709.07124v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-recurrent-nmf-for-speech-separation-by","repo_url":"https://github.com/stwisdom/dr-nmf","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"speech-separation","task_name":"Speech Separation"}],"methods":[{"method_slug":"interpretability","method_name":"Interpretability"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1709.07124","atlas_url":"https://app.syntology.ai/?focus=1709.07124","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}