Papers › Continual self-training with bootstrapped remixing for speech enhancement

Continual self-training with bootstrapped remixing for speech enhancement

19 Oct 2021arXiv:2110.10103archive 2025-07-28

Efthymios Tzinis, Yossi Adi, Vamsi K. Ithapu, Buye Xu, Anurag Kumar

We propose RemixIT, a simple and novel self-supervised training method for speech enhancement. The proposed method is based on a continuously self-training scheme that overcomes limitations from previous studies including assumptions for the in-domain noise distribution and having access to clean target signals. Specifically, a separation teacher model is pre-trained on an out-of-domain dataset and is used to infer estimated target signals for a batch of in-domain mixtures. Next, we bootstrap the mixing process by generating artificial mixtures using permuted estimated clean and noise signals. Finally, the student model is trained using the permuted estimated sources as targets while we periodically update teacher's weights using the latest student model. Our experiments show that RemixIT outperforms several previous state-of-the-art self-supervised methods under multiple speech enhancement tasks. Additionally, RemixIT provides a seamless alternative for semi-supervised and unsupervised domain adaptation for speech enhancement tasks, while being general enough to be applied to any separation task and paired with any separation model.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Domain AdaptationSpeech EnhancementUnsupervised Domain Adaptation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Enhancement Deep Noise Suppression (DNS) Challenge Sudo rm-rf (U=8) PESQ-WB 2.69 #26 of 36 Archive leaderboard report
Speech Enhancement Deep Noise Suppression (DNS) Challenge Sudo rm-rf (U=8) SI-SDR-WB 18.6 #26 of 36 Archive leaderboard report
Speech Enhancement Deep Noise Suppression (DNS) Challenge RemixIT (w Sudo U=32) PESQ-WB 2.60 #28 of 36 Archive leaderboard report
Speech Enhancement Deep Noise Suppression (DNS) Challenge RemixIT (w Sudo U=32) SI-SDR-WB 18.0 #28 of 36 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections