{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/a-unified-neural-architecture-for","title":"A Unified Neural Architecture for Instrumental Audio Tasks","arxiv_id":"1903.00142","date":"2019-03-01","proceeding":null,"authors":["Steven Spratley","Daniel Beck","Trevor Cohn"],"abstract":"Within Music Information Retrieval (MIR), prominent tasks -- including\npitch-tracking, source-separation, super-resolution, and synthesis -- typically\ncall for specialised methods, despite their similarities. Conditional\nGenerative Adversarial Networks (cGANs) have been shown to be highly versatile\nin learning general image-to-image translations, but have not yet been adapted\nacross MIR. In this work, we present an end-to-end supervisable architecture to\nperform all aforementioned audio tasks, consisting of a WaveNet synthesiser\nconditioned on the output of a jointly-trained cGAN spectrogram translator. In\ndoing so, we demonstrate the potential of such flexible techniques to unify MIR\ntasks, promote efficient transfer learning, and converge research to the\nimprovement of powerful, general methods. Finally, to the best of our\nknowledge, we present the first application of GANs to guided instrument\nsynthesis.","url_abs":"http://arxiv.org/abs/1903.00142v1","url_pdf":"http://arxiv.org/pdf/1903.00142v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"a-unified-neural-architecture-for","repo_url":"https://github.com/r9y9/wavenet_vocoder","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}}],"tasks":[{"task_slug":"information-retrieval","task_name":"Information Retrieval"},{"task_slug":"music-information-retrieval","task_name":"Music Information Retrieval"},{"task_slug":"retrieval","task_name":"Retrieval"},{"task_slug":"super-resolution","task_name":"Super-Resolution"},{"task_slug":"transfer-learning","task_name":"Transfer Learning"}],"methods":[{"method_slug":"dilated-causal-convolution","method_name":"Dilated Causal Convolution"},{"method_slug":"mixture-of-logistic-distributions","method_name":"Mixture of Logistic Distributions"},{"method_slug":"wavenet","method_name":"WaveNet"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}