{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/improved-speech-separation-with-time-and","title":"Improved Speech Separation with Time-and-Frequency Cross-domain Joint Embedding and Clustering","arxiv_id":"1904.07845","date":"2019-04-16","proceeding":null,"authors":["Gene-Ping Yang","Chao-I Tuan","Hung-Yi Lee","Lin-shan Lee"],"abstract":"Speech separation has been very successful with deep learning techniques.\nSubstantial effort has been reported based on approaches over spectrogram,\nwhich is well known as the standard time-and-frequency cross-domain\nrepresentation for speech signals. It is highly correlated to the phonetic\nstructure of speech, or \"how the speech sounds\" when perceived by human, but\nprimarily frequency domain features carrying temporal behaviour. Very\nimpressive work achieving speech separation over time domain was reported\nrecently, probably because waveforms in time domain may describe the different\nrealizations of speech in a more precise way than spectrogram. In this paper,\nwe propose a framework properly integrating the above two directions, hoping to\nachieve both purposes. We construct a time-and-frequency feature map by\nconcatenating the 1-dim convolution encoded feature map (for time domain) and\nthe spectrogram (for frequency domain), which was then processed by an\nembedding network and clustering approaches very similar to those used in time\nand frequency domain prior works. In this way, the information in the time and\nfrequency domains, as well as the interactions between them, can be jointly\nconsidered during embedding and clustering. Very encouraging results\n(state-of-the-art to our knowledge) were obtained with WSJ0-2mix dataset in\npreliminary experiments.","url_abs":"http://arxiv.org/abs/1904.07845v1","url_pdf":"http://arxiv.org/pdf/1904.07845v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"improved-speech-separation-with-time-and","repo_url":"https://github.com/r06944010/improved-speech-separation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"clustering","task_name":"Clustering"},{"task_slug":"speech-separation","task_name":"Speech Separation"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-separation-on-wsj0-2mix","task":"Speech Separation","dataset":"WSJ0-2mix","model":"Hybrid-Tasnet","rank_in_archive_order":33,"of":40,"metrics":{"SI-SDRi":"16.6"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}