{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/generalization-challenges-for-neural","title":"Generalization Challenges for Neural Architectures in Audio Source Separation","arxiv_id":"1803.08629","date":"2018-03-23","proceeding":null,"authors":["Shariq Mobin","Brian Cheung","Bruno Olshausen"],"abstract":"Recent work has shown that recurrent neural networks can be trained to\nseparate individual speakers in a sound mixture with high fidelity. Here we\nexplore convolutional neural network models as an alternative and show that\nthey achieve state-of-the-art results with an order of magnitude fewer\nparameters. We also characterize and compare the robustness and ability of\nthese different approaches to generalize under three different test conditions:\nlonger time sequences, the addition of intermittent noise, and different\ndatasets not seen during training. For the last condition, we create a new\ndataset, RealTalkLibri, to test source separation in real-world environments.\nWe show that the acoustics of the environment have significant impact on the\nstructure of the waveform and the overall performance of neural network models,\nwith the convolutional model showing superior ability to generalize to new\nenvironments. The code for our study is available at\nhttps://github.com/ShariqM/source_separation.","url_abs":"http://arxiv.org/abs/1803.08629v2","url_pdf":"http://arxiv.org/pdf/1803.08629v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"generalization-challenges-for-neural","repo_url":"https://github.com/ShariqM/source_separation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"audio-source-separation","task_name":"Audio Source Separation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}