Papers › Multi-View Attention Transfer for Efficient Speech Enhancement

Multi-View Attention Transfer for Efficient Speech Enhancement

22 Aug 2022arXiv:2208.10367archive 2025-07-28

WooSeok Shin, Hyun Joon Park, Jin Sob Kim, Byung Hoon Lee, Sung Won Han

Recent deep learning models have achieved high performance in speech enhancement; however, it is still challenging to obtain a fast and low-complexity model without significant performance degradation. Previous knowledge distillation studies on speech enhancement could not solve this problem because their output distillation methods do not fit the speech enhancement task in some aspects. In this study, we propose multi-view attention transfer (MV-AT), a feature-based distillation, to obtain efficient speech enhancement models in the time domain. Based on the multi-view features extraction model, MV-AT transfers multi-view knowledge of the teacher network to the student network without additional parameters. The experimental results show that the proposed method consistently improved the performance of student models of various sizes on the Valentini and deep noise suppression (DNS) datasets. MANNER-S-8.1GF with our proposed method, a lightweight model for efficient deployment, achieved 15.4x and 4.71x fewer parameters and floating-point operations (FLOPs), respectively, compared to the baseline model with similar performance.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Knowledge DistillationSpeech Enhancement

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Enhancement VoiceBank + DEMAND MANNER-S + MV-AT (8.1GF) CBAK 3.61 #27 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND MANNER-S + MV-AT (8.1GF) COVL 3.82 #27 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND MANNER-S + MV-AT (8.1GF) CSIG 4.45 #27 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND MANNER-S + MV-AT (8.1GF) PESQ (wb) 3.12 #27 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND MANNER-S + MV-AT (8.1GF) Para. (M) 1.38 #27 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND MANNER-S + MV-AT (8.1GF) STOI 95 #27 of 42 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Knowledge Distillation

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections