Home › Datasets › modality › Music

Music datasets

archive 2025-07-28

47 datasets carry the modality tag "Music", ordered by the archive's paper count. Page 1 of 1: 47 shown of 47. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 39 modality tags shown of 39, by dataset count; the full filter by modality, task and language is on /datasets

Music datasets 1–47 of 47

MUSAN is a corpus of music, speech and noise.
204 papers · 0 benchmarks
The MAESTRO dataset contains over 200 hours of paired audio and MIDI recordings from ten years of International Piano-e-Competition.
118 papers · 1 benchmark
MusicCaps is a dataset composed of 5.5k music-text pairs, with rich text descriptions provided by human experts.
84 papers · 1 benchmark
MusicNet is a collection of 330 freely-licensed classical music recordings, together with over 1 million annotated labels indicating the precise time of each note in every recording, the instrument that plays each note, and the note's…
43 papers · 1 benchmark
The MTG-Jamendo dataset is an open dataset for music auto-tagging.
39 papers · 0 benchmarks
The JSB chorales are a set of short, four-voice pieces of music well-noted for their stylistic homogeneity.
33 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
18 papers · 1 benchmark
ASAP (Aligned Scores and Performances)
ASAP is a dataset of 222 digital musical scores aligned with 1068 performances (more than 92 hours) of Western classical piano music.
13 papers · 2 benchmarks
The Lakh Pianoroll Dataset (LPD) is a collection of 174,154 multitrack pianorolls derived from the Lakh MIDI Dataset (LMD).
10 papers · 0 benchmarks
First large-scale symphony generation dataset.
10 papers · 1 benchmark
ATEPP (Automatically Transcribed Expressive Piano Performances)
ATEPP is a dataset of expressive piano performances by virtuoso pianists.
9 papers · 0 benchmarks
AIOZ-GDANCE comprises 16.7 hours of whole-body motion and music audio of group dancing.
7 papers · 1 benchmark
SingFake (SingFake: Singing Voice Deepfake Detection)
The rise of singing voice synthesis presents critical challenges to artists and industry stakeholders over unauthorized voice usage.
7 papers · 0 benchmarks
GTSinger (GTSinger: A Global Multi-Technique Singing Corpus with Realistic Music Scores for All Singing Tasks)
The scarcity of high-quality and multi-task singing datasets significantly hinders the development of diverse controllable and personalized singing tasks, as existing singing datasets suffer from low quality, limited diversity of languages…
6 papers · 0 benchmarks
MuChoMusic is a benchmark designed to evaluate music understanding in multimodal language models focused on audio.
6 papers · 0 benchmarks
The MusicBench dataset is a music audio-text pair dataset that was designed for text-to-music generation purpose and released along with Mustango text-to-music model.
6 papers · 1 benchmark
A large-scale and diverse duet interactive dance dataset.
5 papers · 0 benchmarks
The Song Describer Dataset (SDD) contains ~1.1k captions for 706 permissively licensed music recordings.
5 papers · 1 benchmark
ComMU has 11,144 MIDI samples that consist of short note sequences created by professional composers with their corresponding 12 metadata.
4 papers · 0 benchmarks
For each dataset we provide a short description as well as some characterization metrics.
4 papers · 0 benchmarks
Giantsteps is a dataset that includes songs in major and minor scales for all pitch classes, i.e., a 24-way classification task.
3 papers · 0 benchmarks
A MIDI dataset of 500 4-part chorales generated by the KSChorus algorithm, annotated with results from hundreds of listening test participants, with 500 further unannotated chorales.
3 papers · 0 benchmarks
AIME (AI Music Evaluation Dataset)
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
2 papers · 0 benchmarks
Emomusic (Emotion in Music Database)
1000 songs has been selected from Free Music Archive (FMA).
2 papers · 1 benchmark
Fingerprint Dataset (Neural Audio Fingerprint Dataset)
This dataset includes all music sources, background noises and impulse-reponses (IR) samples and conversation speech that have been used in the work "Neural Audio Fingerprint for High-specific Audio Retrieval based on Contrastive Learning"…
2 papers · 0 benchmarks
IMEMNET (Image-MusicEmotion-Matching-Net)
The Image-MusicEmotion-Matching-Net (IMEMNet) dataset is a dataset for continuous emotion-based image and music matching.
2 papers · 0 benchmarks
Jam-ALT (JamALT: A Formatting-Aware Lyrics Transcription Benchmark)
JamALT is a revision of the JamendoLyrics dataset (80 songs in 4 languages), adapted for use as an automatic lyrics transcription (ALT) benchmark.
2 papers · 5 benchmarks
Lyra Dataset (A Dataset for Greek Traditional and Folk Music)
Lyra is a dataset of 1570 traditional and folk Greek music pieces that includes audio and video (timestamps and links to YouTube videos), along with annotations that describe aspects of particular interest for this dataset, including…
2 papers · 0 benchmarks
The SynthSOD dataset contains more than 47 hours of multitrack music obtained by synthesizing orchestra and ensemble pieces from the Symbolic Orchestral Database (SOD) using Spitfire BBC Symphony Orchestra Professional Library.
2 papers · 0 benchmarks
The YouTube8M-MusicTextClips dataset consists of over 4k high-quality human text descriptions of music found in video clips from the YouTube8M dataset.
2 papers · 0 benchmarks
AVASpeech-SMAD (AVASpeech-SMAD: A Strongly Labelled Speech and Music Activity Detection Dataset with Label Co-Occurrence)
We propose a dataset, AVASpeech-SMAD, to assist speech and music activity detection research.
1 paper · 0 benchmarks
ChMusic is a traditional Chinese music dataset for training model and performance evaluation of musical instrument recognition.
1 paper · 0 benchmarks
We construct a large-scale conducting motion dataset, named ConductorMotion100, by deploying pose estimation on conductor view videos of concert performance recordings collected from online video platforms.
1 paper · 0 benchmarks
Dizi is a dataset of music style of the Northern school and the Southern School.
1 paper · 0 benchmarks
48 multitrack jazz recordings with many annotations.
1 paper · 2 benchmarks
The Haydn Annotation Dataset consists of note onset annotations from 24 experiment participants with varying musical experience.
1 paper · 0 benchmarks
📊 Dataset Details - Name: JamendoMaxCaps - URL: https://huggingface.co/datasets/amaai-lab/JamendoMaxCaps - Content: 362,238 songs with captions generated by Qwen2-Audio Metadata Fields - genre - speed - variable tags 🎯 Rationale 1.
1 paper · 0 benchmarks
This dataset contains annotations for 5000 music files on the following music properties: Melodiousness Articulation Rhythmic stability Rhythmic complexity Dissonance Tonal stability Modality The annotations were given by musicians and…
1 paper · 0 benchmarks
MuVi (MusicVideos)
A dataset of music videos with continuous valence/arousal ratings as well as emotion tags.
1 paper · 0 benchmarks
Nlakh is a dataset for Musical Instrument Retrieval.
1 paper · 0 benchmarks
Virtuoso Strings is a dataset for soft onsets detection for string instruments.
1 paper · 0 benchmarks
XMIDI is a comprehensive, large-scale symbolic music dataset that includes accurate emotion and genre labels, consisting of 108,023 MIDI files.
1 paper · 0 benchmarks
YM2413-MDB is an 80s FM video game music dataset with multi-label emotion annotations.
1 paper · 0 benchmarks
inaGVAD (InaGVAD : a Challenging French TV and Radio Corpus annotated for Voice Activity Detection and Speaker Gender Segmentation)
InaGVAD is a Voice Activity Detection (VAD) and Speaker Gender Segmentation (SGS) dataset designed for representing the acoustic diversity of French TV and Radio programs.
1 paper · 0 benchmarks
jaCappella is a corpus of Japanese a cappella vocal ensembles (jaCappella corpus) for vocal ensemble separation and synthesis.
1 paper · 0 benchmarks
taste-music-dataset (Taste Music Dataset)
This dataset is a patched version of The Taste & Affect Music Database by D.
1 paper · 0 benchmarks
This publicly available data is synthesised audio for woodwind quartets including renderings of each instrument in isolation.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.