Home › Datasets › task › Music Transcription

Music Transcription datasets

archive 2025-07-28

14 datasets carry the task tag "Music Transcription" (the task itself: Music Transcription), ordered by the archive's paper count. Page 1 of 1: 14 shown of 14. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Music Transcription datasets 1–14 of 14

The MAESTRO dataset contains over 200 hours of paired audio and MIDI recordings from ten years of International Piano-e-Competition.
118 papers · 1 benchmark
MusicNet is a collection of 330 freely-licensed classical music recordings, together with over 1 million annotated labels indicating the precise time of each note in every recording, the instrument that plays each note, and the note's…
43 papers · 1 benchmark
Slakh2100 (Synthesized Lakh Dataset)
The Synthesized Lakh (Slakh) Dataset is a dataset for audio source separation that is synthesized from the Lakh MIDI Dataset v0.1 using professional-grade sample-based virtual instruments.
38 papers · 3 benchmarks
URMP (University of Rochester Multi-Modal Musical Performance)
URMP (University of Rochester Multi-Modal Musical Performance) is a dataset for facilitating audio-visual analysis of musical performances.
38 papers · 2 benchmarks
Music21 is an untrimmed video dataset crawled by keyword query from Youtube.
33 papers · 0 benchmarks
ASAP (Aligned Scores and Performances)
ASAP is a dataset of 222 digital musical scores aligned with 1068 performances (more than 92 hours) of Western classical piano music.
13 papers · 2 benchmarks
The CocoChorales Dataset CocoChorales is a dataset consisting of over 1400 hours of audio mixtures containing four-part chorales performed by 13 instruments, all synthesized with realistic-sounding generative models.
7 papers · 0 benchmarks
MAPS (Midi Aligned Piano Dataset)
MAPS – standing for MIDI Aligned Piano Sounds – is a database of MIDI-annotated piano recordings.
7 papers · 1 benchmark
A MIDI dataset of 500 4-part chorales generated by the KSChorus algorithm, annotated with results from hundreds of listening test participants, with 500 further unannotated chorales.
3 papers · 0 benchmarks
This dataset contains transcriptions of the electric guitar performance of 240 tablatures, rendered with different tones.
1 paper · 0 benchmarks
ErhuPT (Erhu Playing Technique Dataset)
This dataset is an audio dataset containing about 1500 audio clips recorded by multiple professional players.
1 paper · 0 benchmarks
Guitar-TECHS (Guitar Tones/Techniques, Excerpts & Chords Dataset)
Guitar-TECHS is a comprehensive dataset featuring a variety of guitar techniques, musical excerpts, chords, and scales.
1 paper · 0 benchmarks
We redistribute a suite of datasets as part of the YourMT3 project.
1 paper · 0 benchmarks
This publicly available data is synthesised audio for woodwind quartets including renderings of each instrument in isolation.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.