{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/sepit-approaching-a-single-channel-speech","title":"SepIt: Approaching a Single Channel Speech Separation Bound","arxiv_id":"2205.11801","date":"2022-05-24","proceeding":null,"authors":["Shahar Lutati","Eliya Nachmani","Lior Wolf"],"abstract":"We present an upper bound for the Single Channel Speech Separation task, which is based on an assumption regarding the nature of short segments of speech. Using the bound, we are able to show that while the recent methods have made significant progress for a few speakers, there is room for improvement for five and ten speakers. We then introduce a Deep neural network, SepIt, that iteratively improves the different speakers' estimation. At test time, SpeIt has a varying number of iterations per test sample, based on a mutual information criterion that arises from our analysis. In an extensive set of experiments, SepIt outperforms the state-of-the-art neural networks for 2, 3, 5, and 10 speakers.","url_abs":"https://arxiv.org/abs/2205.11801v4","url_pdf":"https://arxiv.org/pdf/2205.11801v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"audio-source-separation","task_name":"Audio Source Separation"},{"task_slug":"generalization-bounds","task_name":"Generalization Bounds"},{"task_slug":"multi-speaker-source-separation","task_name":"Multi-Speaker Source Separation"},{"task_slug":"speech-separation","task_name":"Speech Separation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-separation-on-libri10mix","task":"Speech Separation","dataset":"Libri10Mix","model":"SepIt","rank_in_archive_order":2,"of":3,"metrics":{"SI-SDRi":"8.2"},"uses_additional_data":false},{"leaderboard":"/sota/speech-separation-on-libri5mix","task":"Speech Separation","dataset":"Libri5Mix","model":"SepIt","rank_in_archive_order":2,"of":4,"metrics":{"SI-SDRi":"13.7"},"uses_additional_data":true},{"leaderboard":"/sota/speech-separation-on-wsj0-2mix","task":"Speech Separation","dataset":"WSJ0-2mix","model":"SepIt","rank_in_archive_order":14,"of":40,"metrics":{"SI-SDRi":"22.4"},"uses_additional_data":false},{"leaderboard":"/sota/speech-separation-on-wsj0-3mix","task":"Speech Separation","dataset":"WSJ0-3mix","model":"SepIt","rank_in_archive_order":6,"of":9,"metrics":{"SI-SDRi":"20.1"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2205.11801","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}