{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/invariances-and-data-augmentation-for","title":"Invariances and Data Augmentation for Supervised Music Transcription","arxiv_id":"1711.04845","date":"2017-11-13","proceeding":null,"authors":["John Thickstun","Zaid Harchaoui","Dean Foster","Sham M. Kakade"],"abstract":"This paper explores a variety of models for frame-based music transcription,\nwith an emphasis on the methods needed to reach state-of-the-art on human\nrecordings. The translation-invariant network discussed in this paper, which\ncombines a traditional filterbank with a convolutional neural network, was the\ntop-performing model in the 2017 MIREX Multiple Fundamental Frequency\nEstimation evaluation. This class of models shares parameters in the\nlog-frequency domain, which exploits the frequency invariance of music to\nreduce the number of model parameters and avoid overfitting to the training\ndata. All models in this paper were trained with supervision by labeled data\nfrom the MusicNet dataset, augmented by random label-preserving pitch-shift\ntransformations.","url_abs":"http://arxiv.org/abs/1711.04845v1","url_pdf":"http://arxiv.org/pdf/1711.04845v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"invariances-and-data-augmentation-for","repo_url":"https://github.com/jthickstun/thickstun2018invariances","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"music-transcription","task_name":"Music Transcription"},{"task_slug":"translation","task_name":"Translation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}