{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fast-multichannel-source-separation-based-on","title":"Fast Multichannel Source Separation Based on Jointly Diagonalizable Spatial Covariance Matrices","arxiv_id":"1903.03237","date":"2019-03-08","proceeding":"European Association for Signal Processing (EUSIPCO) 2019 9","authors":["Kouhei Sekiguchi","Aditya Arie Nugraha","Yoshiaki Bando","Kazuyoshi Yoshii"],"abstract":"This paper describes a versatile method that accelerates multichannel source\nseparation methods based on full-rank spatial modeling. A popular approach to\nmultichannel source separation is to integrate a spatial model with a source\nmodel for estimating the spatial covariance matrices (SCMs) and power spectral\ndensities (PSDs) of each sound source in the time-frequency domain. One of the\nmost successful examples of this approach is multichannel nonnegative matrix\nfactorization (MNMF) based on a full-rank spatial model and a low-rank source\nmodel. MNMF, however, is computationally expensive and often works poorly due\nto the difficulty of estimating the unconstrained full-rank SCMs. Instead of\nrestricting the SCMs to rank-1 matrices with the severe loss of the spatial\nmodeling ability as in independent low-rank matrix analysis (ILRMA), we\nrestrict the SCMs of each frequency bin to jointly-diagonalizable but still\nfull-rank matrices. For such a fast version of MNMF, we propose a\ncomputationally-efficient and convergence-guaranteed algorithm that is similar\nin form to that of ILRMA. Similarly, we propose a fast version of a\nstate-of-the-art speech enhancement method based on a deep speech model and a\nlow-rank noise model. Experimental results showed that the fast versions of\nMNMF and the deep speech enhancement method were several times faster and\nperformed even better than the original versions of those methods,\nrespectively.","url_abs":"http://arxiv.org/abs/1903.03237v1","url_pdf":"http://arxiv.org/pdf/1903.03237v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"fast-multichannel-source-separation-based-on","repo_url":"https://github.com/sekiguchi92/SpeechEnhancement","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"fast-multichannel-source-separation-based-on","repo_url":"https://github.com/sekiguchi92/eusipco2019","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"fast-multichannel-source-separation-based-on","repo_url":"https://github.com/tky823/ssspy","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}