{"url":"/method/sepformer","slug":"sepformer","name":"SepFormer","full_name":"SepFormer","full_name_withheld":false,"description_markdown":"**SepFormer** is [Transformer](https://paperswithcode.com/methods/category/transformers)-based neural network for speech separation. The SepFormer learns short and long-term dependencies with a multi-scale approach that employs transformers. It is mainly composed of multi-head attention and feed-forward layers. A dual-path framework (introduced by DPRNN) is adopted and [RNNs](https://paperswithcode.com/methods/category/recurrent-neural-networks) are replaced with a multiscale pipeline composed of transformers that learn both short and long-term dependencies. The dual-path framework enables the mitigation of the quadratic complexity of transformers, as transformers in the dual-path framework process smaller chunks.\r\n\r\nThe model is based on the learned-domain masking approach and employs an encoder, a decoder, and a masking network, as shown in the figure. The encoder is fully convolutional, while the decoder employs two Transformers embedded inside the dual-path processing block. The decoder finally reconstructs the separated signals in the time domain by using the masks predicted by the masking network.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Attention is All You Need in Speech Separation","paper":"/paper/attention-is-all-you-need-in-speech","first_author":"Cem Subakan","n_authors":5,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/attention-is-all-you-need-in-speech"},"source":{"url":"https://arxiv.org/abs/2010.13154v2","title":"Attention is All You Need in Speech Separation","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Audio","area_id":"audio","collection":"Speech Separation Models","url":"/methods/category/speech-separation-models","pwc_aliases":["speech-separation"]}],"n_papers_tagged":12,"archive_num_papers":12,"papers_newest_first":[{"paper":"/paper/beyond-speaker-identity-text-guided-target","title":"Beyond Speaker Identity: Text Guided Target Speech Extraction","date":"2025-01-15","arxiv_id":"2501.09169","n_code_links":1,"syntology":null},{"paper":null,"title":"LMAC-TD: Producing Time Domain Explanations for Audio Classifiers","date":"2024-09-13","arxiv_id":"2409.08655","n_code_links":0,"syntology":null},{"paper":"/paper/noise-robust-speech-separation-with-fast","title":"Noise-robust Speech Separation with Fast Generative Correction","date":"2024-06-11","arxiv_id":"2406.07461","n_code_links":1,"syntology":null},{"paper":null,"title":"On Data Sampling Strategies for Training Neural Network Speech Separation Models","date":"2023-04-14","arxiv_id":"2304.07142","n_code_links":0,"syntology":null},{"paper":null,"title":"Towards Real-Time Single-Channel Speech Separation in Noisy and Reverberant Environments","date":"2023-03-14","arxiv_id":"2303.07569","n_code_links":0,"syntology":null},{"paper":null,"title":"X-SepFormer: End-to-end Speaker Extraction Network with Explicit Optimization on Speaker Confusion","date":"2023-03-09","arxiv_id":"2303.05023","n_code_links":0,"syntology":null},{"paper":"/paper/separate-and-diffuse-using-a-pretrained","title":"Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation","date":"2023-01-25","arxiv_id":"2301.10752","n_code_links":0,"syntology":null},{"paper":null,"title":"Improving Target Speaker Extraction with Sparse LDA-transformed Speaker Embeddings","date":"2023-01-16","arxiv_id":"2301.06277","n_code_links":0,"syntology":null},{"paper":null,"title":"Efficient Transformer-based Speech Enhancement Using Long Frames and STFT Magnitudes","date":"2022-06-23","arxiv_id":"2206.11703","n_code_links":0,"syntology":null},{"paper":"/paper/on-using-transformers-for-speech-separation","title":"Exploring Self-Attention Mechanisms for Speech Separation","date":"2022-02-06","arxiv_id":"2202.02884","n_code_links":1,"syntology":null},{"paper":null,"title":"Monaural source separation: From anechoic to reverberant environments","date":"2021-11-15","arxiv_id":"2111.07578","n_code_links":0,"syntology":null},{"paper":"/paper/attention-is-all-you-need-in-speech","title":"Attention is All You Need in Speech Separation","date":"2020-10-25","arxiv_id":"2010.13154","n_code_links":4,"syntology":null}],"papers_shown":12,"tasks":[{"task":"/task/speech-separation","name":"Speech Separation","papers":9},{"task":"/task/decoder","name":"Decoder","papers":2},{"task":"/task/speech-enhancement","name":"Speech Enhancement","papers":2},{"task":"/task/speech-extraction","name":"Speech Extraction","papers":2},{"task":"/task/all","name":"All","papers":1},{"task":"/task/audio-source-separation","name":"Audio Source Separation","papers":1},{"task":"/task/denoising","name":"Denoising","papers":1},{"task":"/task/generalization-bounds","name":"Generalization Bounds","papers":1},{"task":"/task/multi-speaker-source-separation","name":"Multi-Speaker Source Separation","papers":1},{"task":"/task/speaker-verification","name":"Speaker Verification","papers":1},{"task":"/task/target-speaker-extraction","name":"Target Speaker Extraction","papers":1}],"tasks_shown":11,"n_tasks":11,"usage_by_year":[{"year":"2020","papers":1},{"year":"2021","papers":1},{"year":"2022","papers":2},{"year":"2023","papers":5},{"year":"2024","papers":2},{"year":"2025","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/sepformer"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}