{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/tf-attention-net-an-end-to-end-neural-network","title":"Sams-Net: A Sliced Attention-based Neural Network for Music Source Separation","arxiv_id":"1909.05746","date":"2019-09-12","proceeding":null,"authors":["Tingle Li","Jia-Wei Chen","Haowen Hou","Ming Li"],"abstract":"Convolutional Neural Network (CNN) or Long short-term memory (LSTM) based models with the input of spectrogram or waveforms are commonly used for deep learning based audio source separation. In this paper, we propose a Sliced Attention-based neural network (Sams-Net) in the spectrogram domain for the music source separation task. It enables spectral feature interactions with multi-head attention mechanism, achieves easier parallel computing and has a larger receptive field compared with LSTMs and CNNs respectively. Experimental results on the MUSDB18 dataset show that the proposed method, with fewer parameters, outperforms most of the state-of-the-art DNN-based methods.","url_abs":"https://arxiv.org/abs/1909.05746v4","url_pdf":"https://arxiv.org/pdf/1909.05746v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"tf-attention-net-an-end-to-end-neural-network","repo_url":"https://github.com/chenjiawei5/Contribute_Paper","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"audio-source-separation","task_name":"Audio Source Separation"},{"task_slug":"music-source-separation","task_name":"Music Source Separation"}],"methods":[{"method_slug":"attention","method_name":"Attention"},{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"linear-layer","method_name":"Linear Layer"},{"method_slug":"multi-head-attention","method_name":"Multi-Head Attention"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"softmax","method_name":"Softmax"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/music-source-separation-on-musdb18","task":"Music Source Separation","dataset":"MUSDB18","model":"Sams-Net","rank_in_archive_order":23,"of":27,"metrics":{"SDR (avg)":"5.65","SDR (bass)":"5.25","SDR (drums)":"6.63","SDR (other)":"4.09","SDR (vocals)":"6.61"},"uses_additional_data":false}],"syntology":{"syntology_url":null,"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}