{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/on-time-domain-conformer-models-for-monaural","title":"On Time Domain Conformer Models for Monaural Speech Separation in Noisy Reverberant Acoustic Environments","arxiv_id":"2310.06125","date":"2023-10-09","proceeding":null,"authors":["William Ravenscroft","Stefan Goetze","Thomas Hain"],"abstract":"Speech separation remains an important topic for multi-speaker technology researchers. Convolution augmented transformers (conformers) have performed well for many speech processing tasks but have been under-researched for speech separation. Most recent state-of-the-art (SOTA) separation models have been time-domain audio separation networks (TasNets). A number of successful models have made use of dual-path (DP) networks which sequentially process local and global information. Time domain conformers (TD-Conformers) are an analogue of the DP approach in that they also process local and global context sequentially but have a different time complexity function. It is shown that for realistic shorter signal lengths, conformers are more efficient when controlling for feature dimension. Subsampling layers are proposed to further improve computational efficiency. The best TD-Conformer achieves 14.6 dB and 21.2 dB SISDR improvement on the WHAMR and WSJ0-2Mix benchmarks, respectively.","url_abs":"https://arxiv.org/abs/2310.06125v1","url_pdf":"https://arxiv.org/pdf/2310.06125v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"on-time-domain-conformer-models-for-monaural","repo_url":"https://github.com/jwr1995/pubsep","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"computational-efficiency","task_name":"Computational Efficiency"},{"task_slug":"speech-separation","task_name":"Speech Separation"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-separation-on-whamr","task":"Speech Separation","dataset":"WHAMR!","model":"TD-Conformer (XL) + DM","rank_in_archive_order":6,"of":18,"metrics":{"SI-SDRi":"14.6"},"uses_additional_data":false},{"leaderboard":"/sota/speech-separation-on-whamr","task":"Speech Separation","dataset":"WHAMR!","model":"TD-Conformer (L) + DM","rank_in_archive_order":8,"of":18,"metrics":{"SI-SDRi":"13.4"},"uses_additional_data":false},{"leaderboard":"/sota/speech-separation-on-whamr","task":"Speech Separation","dataset":"WHAMR!","model":"TD-Confomer (M) + DM","rank_in_archive_order":14,"of":18,"metrics":{"SI-SDRi":"12"},"uses_additional_data":false},{"leaderboard":"/sota/speech-separation-on-whamr","task":"Speech Separation","dataset":"WHAMR!","model":"TD-Confomer (S)","rank_in_archive_order":16,"of":18,"metrics":{"SI-SDRi":"10.5"},"uses_additional_data":false},{"leaderboard":"/sota/speech-separation-on-wsj0-2mix","task":"Speech Separation","dataset":"WSJ0-2mix","model":"TD-Conformer (XL) + DM","rank_in_archive_order":21,"of":40,"metrics":{"SI-SDRi":"21.2"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}