{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/singing-voice-separation-using-a-deep","title":"Singing Voice Separation Using a Deep Convolutional Neural Network Trained by Ideal Binary Mask and Cross Entropy","arxiv_id":"1812.01278","date":"2018-12-04","proceeding":null,"authors":["Kin Wah Edward Lin","Balamurali B. T.","Enyan Koh","Simon Lui","Dorien Herremans"],"abstract":"Separating a singing voice from its music accompaniment remains an important\nchallenge in the field of music information retrieval. We present a unique\nneural network approach inspired by a technique that has revolutionized the\nfield of vision: pixel-wise image classification, which we combine with cross\nentropy loss and pretraining of the CNN as an autoencoder on singing voice\nspectrograms. The pixel-wise classification technique directly estimates the\nsound source label for each time-frequency (T-F) bin in our spectrogram image,\nthus eliminating common pre- and postprocessing tasks. The proposed network is\ntrained by using the Ideal Binary Mask (IBM) as the target output label. The\nIBM identifies the dominant sound source in each T-F bin of the magnitude\nspectrogram of a mixture signal, by considering each T-F bin as a pixel with a\nmulti-label (for each sound source). Cross entropy is used as the training\nobjective, so as to minimize the average probability error between the target\nand predicted label for each pixel. By treating the singing voice separation\nproblem as a pixel-wise classification task, we additionally eliminate one of\nthe commonly used, yet not easy to comprehend, postprocessing steps: the Wiener\nfilter postprocessing.\n  The proposed CNN outperforms the first runner up in the Music Information\nRetrieval Evaluation eXchange (MIREX) 2016 and the winner of MIREX 2014 with a\ngain of 2.2702 ~ 5.9563 dB global normalized source to distortion ratio (GNSDR)\nwhen applied to the iKala dataset. An experiment with the DSD100 dataset on the\nfull-tracks song evaluation task also shows that our model is able to compete\nwith cutting-edge singing voice separation systems which use multi-channel\nmodeling, data augmentation, and model blending.","url_abs":"http://arxiv.org/abs/1812.01278v1","url_pdf":"http://arxiv.org/pdf/1812.01278v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"singing-voice-separation-using-a-deep","repo_url":"https://github.com/EdwardLin2014/CNN-with-IBM-for-Singing-Voice-Separation","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"singing-voice-separation-using-a-deep","repo_url":"https://github.com/morehovschi/drumsep","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"image-classification","task_name":"Image Classification"},{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"music-information-retrieval","task_name":"Music Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"image-classification","task_name":"image-classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}