Papers › Music Source Separation in the Waveform Domain

Music Source Separation in the Waveform Domain

27 Nov 2019arXiv:1911.13254archive 2025-07-28

Alexandre Défossez, Nicolas Usunier, Léon Bottou, Francis Bach

Source separation for music is the task of isolating contributions, or stems, from different instruments recorded individually and arranged together to form a song. Such components include voice, bass, drums and any other accompaniments.Contrarily to many audio synthesis tasks where the best performances are achieved by models that directly generate the waveform, the state-of-the-art in source separation for music is to compute masks on the magnitude spectrum. In this paper, we compare two waveform domain architectures. We first adapt Conv-Tasnet, initially developed for speech source separation,to the task of music source separation. While Conv-Tasnet beats many existing spectrogram-domain methods, it suffersfrom significant artifacts, as shown by human evaluations. We propose instead Demucs, a novel waveform-to-waveform model,with a U-Net structure and bidirectional LSTM.Experiments on the MusDB dataset show that, with proper data augmentation, Demucs beats allexisting state-of-the-art architectures, including Conv-Tasnet, with 6.3 SDR on average, (and up to 6.8 with 150 extra training songs, even surpassing the IRM oracle for the bass source).Using recent development in model quantization, Demucs can be compressed down to 120MBwithout any loss of accuracy.We also provide human evaluations, showing that Demucs benefit from a large advantagein terms of the naturalness of the audio. However, it suffers from some bleeding,especially between the vocals and other source.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

facebookresearch/demucs officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Audio GenerationAudio SynthesisData AugmentationMulti-task Audio Source SeperationMusic Source SeparationQuantization

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Music Source Separation MUSDB18 DEMUCS (extra) SDR (avg) 6.79 #10 of 27 Archive leaderboard report
Music Source Separation MUSDB18 DEMUCS (extra) SDR (bass) 7.60 #10 of 27 Archive leaderboard report
Music Source Separation MUSDB18 DEMUCS (extra) SDR (drums) 7.58 #10 of 27 Archive leaderboard report
Music Source Separation MUSDB18 DEMUCS (extra) SDR (other) 4.69 #10 of 27 Archive leaderboard report
Music Source Separation MUSDB18 DEMUCS (extra) SDR (vocals) 7.29 #10 of 27 Archive leaderboard report
Music Source Separation MUSDB18 DEMUCS SDR (avg) 6.28 #16 of 27 Archive leaderboard report
Music Source Separation MUSDB18 DEMUCS SDR (bass) 7.01 #16 of 27 Archive leaderboard report
Music Source Separation MUSDB18 DEMUCS SDR (drums) 6.86 #16 of 27 Archive leaderboard report
Music Source Separation MUSDB18 DEMUCS SDR (other) 4.42 #16 of 27 Archive leaderboard report
Music Source Separation MUSDB18 DEMUCS SDR (vocals) 6.84 #16 of 27 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections