{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/joint-optimization-of-masks-and-deep","title":"Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source Separation","arxiv_id":"1502.04149","date":"2015-02-13","proceeding":null,"authors":["Po-Sen Huang","Minje Kim","Mark Hasegawa-Johnson","Paris Smaragdis"],"abstract":"Monaural source separation is important for many real world applications. It\nis challenging because, with only a single channel of information available,\nwithout any constraints, an infinite number of solutions are possible. In this\npaper, we explore joint optimization of masking functions and deep recurrent\nneural networks for monaural source separation tasks, including monaural speech\nseparation, monaural singing voice separation, and speech denoising. The joint\noptimization of the deep recurrent neural networks with an extra masking layer\nenforces a reconstruction constraint. Moreover, we explore a discriminative\ncriterion for training neural networks to further enhance the separation\nperformance. We evaluate the proposed system on the TSP, MIR-1K, and TIMIT\ndatasets for speech separation, singing voice separation, and speech denoising\ntasks, respectively. Our approaches achieve 2.30--4.98 dB SDR gain compared to\nNMF models in the speech separation task, 2.30--2.48 dB GNSDR gain and\n4.32--5.42 dB GSIR gain compared to existing models in the singing voice\nseparation task, and outperform NMF and DNN baselines in the speech denoising\ntask.","url_abs":"http://arxiv.org/abs/1502.04149v4","url_pdf":"http://arxiv.org/pdf/1502.04149v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"joint-optimization-of-masks-and-deep","repo_url":"https://github.com/bill9800/speech_separation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"joint-optimization-of-masks-and-deep","repo_url":"https://github.com/vitrioil/Speech-Separation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"speech-denoising","task_name":"Speech Denoising"},{"task_slug":"speech-separation","task_name":"Speech Separation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1502.04149","atlas_url":"https://app.syntology.ai/?focus=1502.04149","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}