{"url":"/sota/speech-separation-on-libri2mix","task":{"name":"Speech Separation","url":"/task/speech-separation","note":null},"dataset":{"name":"Libri2Mix","url":"/dataset/librimix"},"category":"Speech","categories":["Speech"],"category_note":null,"description":"The task of extracting all overlapping speech sources in a given mixed speech signal refers to the **Speech Separation**. Speech Separation is a special scenario of source separation problem, where the focus is only on the overlapping speech signal sources and other interferences such as music or noise signals are not the main concern of the study. A recent representative Github project can be referred to [ClearerVoice-Studio](https://github.com/modelscope/ClearerVoice-Studio).\r\n\r\n\r\n<span class=\"description-source\">Source: [A Unified Framework for Speech Separation ](https://arxiv.org/abs/1912.07814)</span>\r\n\r\nImage credit: [Speech Separation of A Target Speaker Based on Deep Neural Networks](http://staff.ustc.edu.cn/~jundu/Publications/publications/ICSP2014_Du.pdf)","description_from":"task","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","rank":"the archive's row order at snapshot; not re-ranked","rows_end_at":"2025-07-28","rows_withheld_as_spam":0,"metric_values":"the archive's strings, untouched"},"metrics":["SI-SDRi","SDRi","Number of parameters (M)","SDR"],"metric_direction":{"note":"inferred from the metric name only (the archive records no direction); null = not inferred, chart draws points only","by_metric":{"SI-SDRi":null,"SDRi":null,"Number of parameters (M)":null,"SDR":null}},"counts":{"rows":10,"rows_with_code":8,"rows_with_paper_page":10,"rows_dated":10,"rows_using_additional_data":2},"rows":[{"rank_in_archive_order":1,"model":"MossFormer2 (w speed perturb)","metrics":{"SI-SDRi":"22.2"},"uses_additional_data":false,"paper_date":"2023-12-19","paper":"/paper/mossformer2-combining-transformer-and-rnn-1","paper_url":"https://arxiv.org/abs/2312.11825v2","paper_title":"MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation","code":"https://github.com/modelscope/ClearerVoice-Studio","n_code_links":2,"syntology":{"n_ran":4,"n_unverified":0,"n_samples":4,"n_pointer_only_licence":0}},{"rank_in_archive_order":2,"model":"TF-Locoformer (M)","metrics":{"Number of parameters (M)":"15","SDRi":"22.2","SI-SDRi":"22.1"},"uses_additional_data":false,"paper_date":"2024-08-06","paper":"/paper/tf-locoformer-transformer-with-local-modeling","paper_url":"https://arxiv.org/abs/2408.03440v1","paper_title":"TF-Locoformer: Transformer with Local Modeling by Convolution for Speech Separation and Enhancement","code":"https://github.com/merlresearch/tf-locoformer","n_code_links":1,"syntology":null},{"rank_in_archive_order":3,"model":"MossFormer2 (w/o DM)","metrics":{"SI-SDRi":"21.7"},"uses_additional_data":false,"paper_date":"2023-12-19","paper":"/paper/mossformer2-combining-transformer-and-rnn-1","paper_url":"https://arxiv.org/abs/2312.11825v2","paper_title":"MossFormer2: Combining Transformer and RNN-Free Recurrent Network for Enhanced Time-Domain Monaural Speech Separation","code":"https://github.com/modelscope/ClearerVoice-Studio","n_code_links":2,"syntology":{"n_ran":4,"n_unverified":0,"n_samples":4,"n_pointer_only_licence":0}},{"rank_in_archive_order":4,"model":"Separate And Diffuse","metrics":{"SI-SDRi":"21.5"},"uses_additional_data":false,"paper_date":"2023-01-25","paper":"/paper/separate-and-diffuse-using-a-pretrained","paper_url":"https://arxiv.org/abs/2301.10752v2","paper_title":"Separate And Diffuse: Using a Pretrained Diffusion Model for Improving Source Separation","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":5,"model":"WHYV","metrics":{"SDR":"17.2458","SI-SDRi":"17.5"},"uses_additional_data":false,"paper_date":"2024-10-01","paper":"/paper/wanna-hear-your-voice-adaptive-effective-and","paper_url":"https://arxiv.org/abs/2410.00527v4","paper_title":"Wanna hear your voice? A sample is all we need!","code":null,"n_code_links":0,"syntology":null},{"rank_in_archive_order":6,"model":"TDANet Large","metrics":{"SI-SDRi":"17.4"},"uses_additional_data":false,"paper_date":"2022-09-30","paper":"/paper/an-efficient-encoder-decoder-architecture","paper_url":"https://arxiv.org/abs/2209.15200v5","paper_title":"An efficient encoder-decoder architecture with top-down attention for speech separation","code":"https://github.com/JusperLee/TDANet","n_code_links":1,"syntology":{"n_ran":13,"n_unverified":5,"n_samples":18,"n_pointer_only_licence":0}},{"rank_in_archive_order":7,"model":"TDANet","metrics":{"SI-SDRi":"16.9"},"uses_additional_data":false,"paper_date":"2022-09-30","paper":"/paper/an-efficient-encoder-decoder-architecture","paper_url":"https://arxiv.org/abs/2209.15200v5","paper_title":"An efficient encoder-decoder architecture with top-down attention for speech separation","code":"https://github.com/JusperLee/TDANet","n_code_links":1,"syntology":{"n_ran":13,"n_unverified":5,"n_samples":18,"n_pointer_only_licence":0}},{"rank_in_archive_order":8,"model":"Conv-Tasnet (Libri1Mix speech enhancement pre-trained)","metrics":{"SDRi":"14.6","SI-SDRi":"14.1"},"uses_additional_data":true,"paper_date":"2020-10-29","paper":"/paper/self-supervised-pre-training-reduces-label","paper_url":"https://arxiv.org/abs/2010.15366v3","paper_title":"Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training","code":"https://github.com/SungFeng-Huang/SSL-pretraining-separation","n_code_links":1,"syntology":null},{"rank_in_archive_order":9,"model":"Conv-Tasnet (Libri1Mix speech enhancement multi-task)","metrics":{"SDRi":"14.1","SI-SDRi":"13.7"},"uses_additional_data":true,"paper_date":"2020-10-29","paper":"/paper/self-supervised-pre-training-reduces-label","paper_url":"https://arxiv.org/abs/2010.15366v3","paper_title":"Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training","code":"https://github.com/SungFeng-Huang/SSL-pretraining-separation","n_code_links":1,"syntology":null},{"rank_in_archive_order":10,"model":"Conv-Tasnet","metrics":{"SDRi":"13.6","SI-SDRi":"13.2"},"uses_additional_data":false,"paper_date":"2020-10-29","paper":"/paper/self-supervised-pre-training-reduces-label","paper_url":"https://arxiv.org/abs/2010.15366v3","paper_title":"Stabilizing Label Assignment for Speech Separation by Self-supervised Pre-training","code":"https://github.com/SungFeng-Huang/SSL-pretraining-separation","n_code_links":1,"syntology":null}],"since_archive":{"present":false,"note":"No Syntology-extracted rows are published in this build."},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per row: N of M harvested code samples from that row's paper executed on a synthesized fixture; the other M-N are unverified. Not a reproduction of the row's number; not a correctness claim. n_pointer_only_licence counts samples the site points at rather than redistributes (a licence axis, independent of ran/unverified).","rows_with_graph_line":4,"rows_with_any_sample_ran":4,"distinct_papers_with_graph_line":2,"distinct_papers_with_any_sample_ran":2,"samples_over_distinct_papers":{"n_ran":17,"n_unverified":5,"n_samples":22,"n_pointer_only_licence":0,"note":"each paper (arXiv id) counted once, however many rows it is behind; this is the page-level figure"},"samples_row_weighted":{"n_ran":34,"n_unverified":10,"n_samples":44,"n_pointer_only_licence":0,"note":"row-weighted: a paper behind several rows is counted once per row; inflated relative to samples_over_distinct_papers by design, kept for readers summing the per-row syntology blocks"}}}