{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-using-transformers-for-speech-separation","title":"Exploring Self-Attention Mechanisms for Speech Separation","arxiv_id":"2202.02884","date":"2022-02-06","proceeding":null,"authors":["Cem Subakan","Mirco Ravanelli","Samuele Cornell","Francois Grondin","Mirko Bronzi"],"abstract":"Transformers have enabled impressive improvements in deep learning. They often outperform recurrent and convolutional models in many tasks while taking advantage of parallel processing. Recently, we proposed the SepFormer, which obtains state-of-the-art performance in speech separation with the WSJ0-2/3 Mix datasets. This paper studies in-depth Transformers for speech separation. In particular, we extend our previous findings on the SepFormer by providing results on more challenging noisy and noisy-reverberant datasets, such as LibriMix, WHAM!, and WHAMR!. Moreover, we extend our model to perform speech enhancement and provide experimental evidence on denoising and dereverberation tasks. Finally, we investigate, for the first time in speech separation, the use of efficient self-attention mechanisms such as Linformers, Lonformers, and ReFormers. We found that they reduce memory requirements significantly. For example, we show that the Reformer-based attention outperforms the popular Conv-TasNet model on the WSJ0-2Mix dataset while being faster at inference and comparable in terms of memory consumption.","url_abs":"https://arxiv.org/abs/2202.02884v2","url_pdf":"https://arxiv.org/pdf/2202.02884v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-using-transformers-for-speech-separation","repo_url":"https://github.com/speechbrain/speechbrain","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"denoising","task_name":"Denoising"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"},{"task_slug":"speech-separation","task_name":"Speech Separation"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"convtasnet","method_name":"ConvTasNet"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"layer-normalization","method_name":"Layer Normalization"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"prelu","method_name":"PReLU"},{"method_slug":"position-wise-feed-forward-layer","method_name":"Position-Wise Feed-Forward Layer"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"sepformer","method_name":"SepFormer"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-enhancement-on-wham","task":"Speech Enhancement","dataset":"WHAM!","model":"SepFormer","rank_in_archive_order":1,"of":1,"metrics":{"PESQ":"3.07","SDR":"15.04","SI-SNR":"14.35"},"uses_additional_data":false},{"leaderboard":"/sota/speech-enhancement-on-whamr","task":"Speech Enhancement","dataset":"WHAMR!","model":"SepFormer","rank_in_archive_order":1,"of":4,"metrics":{"PESQ":"2.84","SDR":"12.29","SI-SNR":"10.58"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2202.02884","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}