Papers › Utterance Weighted Multi-Dilation Temporal Convolutional Networks for Monaural Speech...

Utterance Weighted Multi-Dilation Temporal Convolutional Networks for Monaural Speech Dereverberation

17 May 2022arXiv:2205.08455archive 2025-07-28

William Ravenscroft, Stefan Goetze, Thomas Hain

Speech dereverberation is an important stage in many speech technology applications. Recent work in this area has been dominated by deep neural network models. Temporal convolutional networks (TCNs) are deep learning models that have been proposed for sequence modelling in the task of dereverberating speech. In this work a weighted multi-dilation depthwise-separable convolution is proposed to replace standard depthwise-separable convolutions in TCN models. This proposed convolution enables the TCN to dynamically focus on more or less local information in its receptive field at each convolutional block in the network. It is shown that this weighted multi-dilation temporal convolutional network (WD-TCN) consistently outperforms the TCN across various model configurations and using the WD-TCN model is a more parameter efficient method to improve the performance of the model than increasing the number of convolutional blocks. The best performance improvement over the baseline TCN is 0.55 dB scale-invariant signal-to-distortion ratio (SISDR) and the best performing WD-TCN model attains 12.26 dB SISDR on the WHAMR dataset.

PaperPDFCode

Code

jwr1995/wd-tcn officialmentioned in papermentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Speech Dereverberation

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Dereverberation WHAMR! WD-TCN ESTOI 93.5 #1 of 3 Archive leaderboard report
Speech Dereverberation WHAMR! WD-TCN PESQ 3.5 #1 of 3 Archive leaderboard report
Speech Dereverberation WHAMR! WD-TCN SI-SDR 12.26 #1 of 3 Archive leaderboard report
Speech Dereverberation WHAMR! WD-TCN SRMR 8.8 #1 of 3 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Convolution

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections