{"url":"/dataset/bach-doodle","name":"Bach Doodle","full_name":"Bach Doodle","description_markdown":"The **Bach Doodle** Dataset is composed of 21.6 million harmonizations submitted from the Bach Doodle. The dataset contains both metadata about the composition (such as the country of origin and feedback), as well as a MIDI of the user-entered melody and a MIDI of the generated harmonization. The dataset contains about 6 years of user entered music.\n\nSource: [https://magenta.tensorflow.org/datasets/bach-doodle](https://magenta.tensorflow.org/datasets/bach-doodle)\nImage Source: [https://magenta.tensorflow.org/datasets/bach-doodle](https://magenta.tensorflow.org/datasets/bach-doodle)","description_withheld":null,"homepage":"https://magenta.tensorflow.org/datasets/bach-doodle","introduced_date":"2019-01-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/the-bach-doodle-approachable-music","title":"The Bach Doodle: Approachable music composition with machine learning at scale","first_author":"Cheng-Zhi Anna Huang","url":null},"license":null,"modalities":[{"name":"Audio","url":"/datasets/modality/audio"}],"tasks":[{"name":"Information Retrieval","url":"/task/information-retrieval","datasets_with_task":"/datasets/task/information-retrieval"},{"name":"Quantization","url":"/task/quantization","datasets_with_task":"/datasets/task/quantization"},{"name":"Music Generation","url":"/task/music-generation","datasets_with_task":"/datasets/task/music-generation"}],"languages":[],"variants":["Bach Doodle"],"data_loaders":[],"num_papers_in_archive":4,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}