Home › Datasets › task › Music Source Separation

Music Source Separation datasets

archive 2025-07-28

9 datasets carry the task tag "Music Source Separation" (the task itself: Music Source Separation), ordered by the archive's paper count. Page 1 of 1: 9 shown of 9. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Music Source Separation datasets 1–9 of 9

The MUSDB18 is a dataset of 150 full lengths music tracks (~10h duration) of different genres along with their isolated drums, bass, vocals and others stems.
106 papers · 2 benchmarks
MedleyDB, is a dataset of annotated, royalty-free multitrack recordings.
47 papers · 0 benchmarks
Slakh2100 (Synthesized Lakh Dataset)
The Synthesized Lakh (Slakh) Dataset is a dataset for audio source separation that is synthesized from the Lakh MIDI Dataset v0.1 using professional-grade sample-based virtual instruments.
38 papers · 3 benchmarks
MIR-1K (Multimedia Information Retrieval lab, 1000 song clips) is a dataset designed for singing voice separation.
21 papers · 0 benchmarks
MUSDB18-HQ is a high-quality version of the MUSDB18 music tracks dataset.
15 papers · 1 benchmark
The CocoChorales Dataset CocoChorales is a dataset consisting of over 1400 hours of audio mixtures containing four-part chorales performed by 13 instruments, all synthesized with realistic-sounding generative models.
7 papers · 0 benchmarks
The MuseScore dataset is a collection of 344,166 audio and MIDI pairs downloaded from MuseScore website.
2 papers · 0 benchmarks
The SynthSOD dataset contains more than 47 hours of multitrack music obtained by synthesizing orchestra and ensemble pieces from the Symbolic Orchestral Database (SOD) using Spitfire BBC Symphony Orchestra Professional Library.
2 papers · 0 benchmarks
This publicly available data is synthesised audio for woodwind quartets including renderings of each instrument in isolation.
0 papers · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.