{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/remixit-continual-self-training-of-speech","title":"RemixIT: Continual self-training of speech enhancement models via bootstrapped remixing","arxiv_id":"2202.08862","date":"2022-02-17","proceeding":null,"authors":["Efthymios Tzinis","Yossi Adi","Vamsi Krishna Ithapu","Buye Xu","Paris Smaragdis","Anurag Kumar"],"abstract":"We present RemixIT, a simple yet effective self-supervised method for training speech enhancement without the need of a single isolated in-domain speech nor a noise waveform. Our approach overcomes limitations of previous methods which make them dependent on clean in-domain target signals and thus, sensitive to any domain mismatch between train and test samples. RemixIT is based on a continuous self-training scheme in which a pre-trained teacher model on out-of-domain data infers estimated pseudo-target signals for in-domain mixtures. Then, by permuting the estimated clean and noise signals and remixing them together, we generate a new set of bootstrapped mixtures and corresponding pseudo-targets which are used to train the student network. Vice-versa, the teacher periodically refines its estimates using the updated parameters of the latest student models. Experimental results on multiple speech enhancement datasets and tasks not only show the superiority of our method over prior approaches but also showcase that RemixIT can be combined with any separation model as well as be applied towards any semi-supervised and unsupervised domain adaptation task. Our analysis, paired with empirical evidence, sheds light on the inside functioning of our self-training scheme wherein the student model keeps obtaining better performance while observing severely degraded pseudo-targets.","url_abs":"https://arxiv.org/abs/2202.08862v3","url_pdf":"https://arxiv.org/pdf/2202.08862v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"remixit-continual-self-training-of-speech","repo_url":"https://github.com/etzinis/unsup_speech_enh_adaptation","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"remixit-continual-self-training-of-speech","repo_url":"https://github.com/udase-chime2023/baseline","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"domain-adaptation","task_name":"Domain Adaptation"},{"task_slug":"speech-enhancement","task_name":"Speech Enhancement"},{"task_slug":"unsupervised-domain-adaptation","task_name":"Unsupervised Domain Adaptation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/speech-enhancement-on-deep-noise-suppression","task":"Speech Enhancement","dataset":"Deep Noise Suppression (DNS) Challenge","model":"Sudo rm -rf (U=32)","rank_in_archive_order":17,"of":36,"metrics":{"PESQ-WB":"2.95","SI-SDR-WB":"19.7"},"uses_additional_data":false},{"leaderboard":"/sota/speech-enhancement-on-deep-noise-suppression","task":"Speech Enhancement","dataset":"Deep Noise Suppression (DNS) Challenge","model":"RemixIT (w Sudo U=32)","rank_in_archive_order":29,"of":36,"metrics":{"PESQ-WB":"2.34","SI-SDR-WB":"16.0"},"uses_additional_data":true}],"syntology":{"syntology_url":"https://syntology.ai/paper/2202.08862","atlas_url":"https://app.syntology.ai/?focus=2202.08862","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2202.08862"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/udase-chime2023/baseline","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/etzinis/unsup_speech_enh_adaptation","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":4},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1},"listed":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"59bcee2a1aec4892","entry":"compute_sisdr","repo":"udase-chime2023/baseline","repo_kind":"listed","path":"baseline/metrics/sisdr_metric.py","file_url":"https://github.com/udase-chime2023/baseline/blob/HEAD/baseline/metrics/sisdr_metric.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"59bcee2a1aec4892"}},{"code_sha256_prefix":"b05bd1c49be2b5a5","entry":"normalize_waveform","repo":"etzinis/unsup_speech_enh_adaptation","repo_kind":"official","path":"baseline/run_remixit.py","file_url":"https://github.com/etzinis/unsup_speech_enh_adaptation/blob/HEAD/baseline/run_remixit.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"b05bd1c49be2b5a5"}},{"code_sha256_prefix":"105a42f915805d8c","entry":"report_losses_mean_and_std","repo":"etzinis/unsup_speech_enh_adaptation","repo_kind":"official","path":"baseline/utils/cometml_logger.py","file_url":"https://github.com/etzinis/unsup_speech_enh_adaptation/blob/HEAD/baseline/utils/cometml_logger.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"105a42f915805d8c"}},{"code_sha256_prefix":"6cb3ac65f741ebb5","entry":"tuple_availavle_speech","repo":"etzinis/unsup_speech_enh_adaptation","repo_kind":"official","path":"baseline/utils/cmd_parser.py","file_url":"https://github.com/etzinis/unsup_speech_enh_adaptation/blob/HEAD/baseline/utils/cmd_parser.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6cb3ac65f741ebb5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}