Home › Datasets › task › Music Information Retrieval
Music Information Retrieval datasets
archive 2025-07-28
26 datasets carry the task tag "Music Information Retrieval" (the task itself: Music Information Retrieval), ordered by the archive's paper count. Page 1 of 1: 26 shown of 26. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Music Information Retrieval datasets 1–26 of 26
The Free Music Archive (FMA) is a large-scale dataset for evaluating several tasks in Music Information Retrieval.
128 papers · 2 benchmarks
MedleyDB, is a dataset of annotated, royalty-free multitrack recordings.
47 papers · 0 benchmarks
RWC (Real World Computing Music Database)
The RWC (Real World Computing) Music Database is a copyright-cleared music database (DB) that is available to researchers as a common foundation for research.
43 papers · 0 benchmarks
The MTG-Jamendo dataset is an open dataset for music auto-tagging.
39 papers · 0 benchmarks
URMP (University of Rochester Multi-Modal Musical Performance)
URMP (University of Rochester Multi-Modal Musical Performance) is a dataset for facilitating audio-visual analysis of musical performances.
38 papers · 2 benchmarks
GuitarSet is a dataset of high-quality guitar recordings and rich annotations.
31 papers · 2 benchmarks
EMOPIA (A Multi-Modal Pop Piano Dataset For Emotion Recognition and Emotion-based Music Generation)
EMOPIA (pronounced ‘yee-mò-pi-uh’) dataset is a shared multi-modal (audio and MIDI) database focusing on perceived emotion in pop piano music, to facilitate research on various tasks related to music emotion.
23 papers · 0 benchmarks
The iKala dataset is a singing voice separation dataset that comprises of 252 30-second excerpts sampled from 206 iKala songs (plus 100 hidden excerpts reserved for MIREX data mining contest).
20 papers · 1 benchmark
The Lakh Pianoroll Dataset (LPD) is a collection of 174,154 multitrack pianorolls derived from the Lakh MIDI Dataset (LMD).
10 papers · 0 benchmarks
GiantMIDI-Piano contains 10,854 unique piano solo pieces composed by 2,786 composers.
9 papers · 0 benchmarks
OpenMIC-2018 is an instrument recognition dataset containing 20,000 examples of Creative Commons-licensed music available on the Free Music Archive.
8 papers · 1 benchmark
The YouTube-100M data set consists of 100 million YouTube videos: 70M training videos, 10M evaluation videos, and 20M validation videos.
8 papers · 0 benchmarks
MSSD (Music Streaming Sessions Dataset)
The Spotify Music Streaming Sessions Dataset (MSSD) consists of 160 million streaming sessions with associated user interactions, audio features and metadata describing the tracks streamed during the sessions, and snapshots of the…
7 papers · 1 benchmark
GoodSounds dataset contains around 28 hours of recordings of single notes and scales played by 15 different professional musicians, all of them holding a music degree and having some expertise in teaching.
4 papers · 0 benchmarks
MuMu is a new dataset of more than 31k albums classified into 250 genre classes.
4 papers · 0 benchmarks
jazznet is a dataset of piano patterns for music audio machine learning research.
4 papers · 0 benchmarks
COSIAN (a collection of singing voice annotation)
COSIAN is an annotation collection of Japanese popular (J-POP) songs, focusing on singing style and expression of famous solo-singers.
3 papers · 0 benchmarks
This dataset includes all music sources, background noises and impulse-reponses (IR) samples and conversation speech that have been used in the work "Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning"…
2 papers · 0 benchmarks
CAL10K (Computer Audition Lab 10000)
The CAL10K dataset (introduced as Swat10k) contains 10,870 songs that are weakly-labelled using a tag vocabulary of 475 acoustic tags and 153 genre tags.
1 paper · 0 benchmarks
Guitar-TECHS (Guitar Tones/Techniques, Excerpts & Chords Dataset)
Guitar-TECHS is a comprehensive dataset featuring a variety of guitar techniques, musical excerpts, chords, and scales.
1 paper · 0 benchmarks
The Haydn Annotation Dataset consists of note onset annotations from 24 experiment participants with varying musical experience.
1 paper · 0 benchmarks
Introduction The Niko Chord Progression Dataset is used in AccoMontage2.
1 paper · 0 benchmarks
Nlakh is a dataset for Musical Instrument Retrieval.
1 paper · 0 benchmarks
PIAST (PIAST: A Multimodal Piano Dataset with Audio, Symbolic and Text)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Virtuoso Strings is a dataset for soft onsets detection for string instruments.
1 paper · 0 benchmarks
This publicly available data is synthesised audio for woodwind quartets including renderings of each instrument in isolation.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.