{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/enabling-factorized-piano-music-modeling-and","title":"Enabling Factorized Piano Music Modeling and Generation with the MAESTRO Dataset","arxiv_id":"1810.12247","date":"2018-10-29","proceeding":"ICLR 2019 5","authors":["Curtis Hawthorne","Andriy Stasyuk","Adam Roberts","Ian Simon","Cheng-Zhi Anna Huang","Sander Dieleman","Erich Elsen","Jesse Engel","Douglas Eck"],"abstract":"Generating musical audio directly with neural networks is notoriously\ndifficult because it requires coherently modeling structure at many different\ntimescales. Fortunately, most music is also highly structured and can be\nrepresented as discrete note events played on musical instruments. Herein, we\nshow that by using notes as an intermediate representation, we can train a\nsuite of models capable of transcribing, composing, and synthesizing audio\nwaveforms with coherent musical structure on timescales spanning six orders of\nmagnitude (~0.1 ms to ~100 s), a process we call Wave2Midi2Wave. This large\nadvance in the state of the art is enabled by our release of the new MAESTRO\n(MIDI and Audio Edited for Synchronous TRacks and Organization) dataset,\ncomposed of over 172 hours of virtuosic piano performances captured with fine\nalignment (~3 ms) between note labels and audio waveforms. The networks and the\ndataset together present a promising approach toward creating new expressive\nand interpretable neural models of music.","url_abs":"http://arxiv.org/abs/1810.12247v5","url_pdf":"http://arxiv.org/pdf/1810.12247v5.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"enabling-factorized-piano-music-modeling-and","repo_url":"https://github.com/BShakhovsky/PolyphonicPianoTranscription","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"enabling-factorized-piano-music-modeling-and","repo_url":"https://github.com/Kangmo/kogenta","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}},{"paper_slug":"enabling-factorized-piano-music-modeling-and","repo_url":"https://github.com/benadar293/benadar293.github.io","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"NOASSERTION"}},{"paper_slug":"enabling-factorized-piano-music-modeling-and","repo_url":"https://github.com/junhoyeo/magenta-school-song","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"music-generation","task_name":"Music Generation"},{"task_slug":"music-modeling","task_name":"Music Modeling"},{"task_slug":"music-transcription","task_name":"Music Transcription"},{"task_slug":"piano-music-modeling","task_name":"Piano Music Modeling"}],"methods":[],"datasets_introduced":[{"slug":"maestro","name":"MAESTRO","full_name":"MAESTRO"}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1810.12247","atlas_url":"https://app.syntology.ai/?focus=1810.12247","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}