{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/deep-karaoke-extracting-vocals-from-musical","title":"Deep Karaoke: Extracting Vocals from Musical Mixtures Using a Convolutional Deep Neural Network","arxiv_id":"1504.04658","date":"2015-04-17","proceeding":null,"authors":["Andrew J. R. Simpson","Gerard Roma","Mark D. Plumbley"],"abstract":"Identification and extraction of singing voice from within musical mixtures\nis a key challenge in source separation and machine audition. Recently, deep\nneural networks (DNN) have been used to estimate 'ideal' binary masks for\ncarefully controlled cocktail party speech separation problems. However, it is\nnot yet known whether these methods are capable of generalizing to the\ndiscrimination of voice and non-voice in the context of musical mixtures. Here,\nwe trained a convolutional DNN (of around a billion parameters) to provide\nprobabilistic estimates of the ideal binary mask for separation of vocal sounds\nfrom real-world musical mixtures. We contrast our DNN results with more\ntraditional linear methods. Our approach may be useful for automatic removal of\nvocal sounds from musical mixtures for 'karaoke' type applications.","url_abs":"http://arxiv.org/abs/1504.04658v1","url_pdf":"http://arxiv.org/pdf/1504.04658v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"deep-karaoke-extracting-vocals-from-musical","repo_url":"https://github.com/ishandutta2007/Speech-Denoising-Landscape","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"speech-separation","task_name":"Speech Separation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1504.04658","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}