Papers › DeFT-AN: Dense Frequency-Time Attentive Network for Multichannel Speech Enhancement
DeFT-AN: Dense Frequency-Time Attentive Network for Multichannel Speech Enhancement
Dongheon Lee, Jung-Woo Choi
In this study, we propose a dense frequency-time attentive network (DeFT-AN) for multichannel speech enhancement. DeFT-AN is a mask estimation network that predicts a complex spectral masking pattern for suppressing the noise and reverberation embedded in the short-time Fourier transform (STFT) of an input signal. The proposed mask estimation network incorporates three different types of blocks for aggregating information in the spatial, spectral, and temporal dimensions. It utilizes a spectral transformer with a modified feed-forward network and a temporal conformer with sequential dilated convolutions. The use of dense blocks and transformers dedicated to the three different characteristics of audio signals enables more comprehensive enhancement in noisy and reverberant environments. The remarkable performance of DeFT-AN over state-of-the-art multichannel models is demonstrated based on two popular noisy and reverberant datasets in terms of various metrics for speech quality and intelligibility.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Speech Dereverberation | spatialized WSJCAM0 | DeFT-AN | PESQ | 3.63 | #1 of 1 | Archive leaderboard | report |
| Speech Dereverberation | spatialized WSJCAM0 | DeFT-AN | SI-SDR | 15.7 | #1 of 1 | Archive leaderboard | report |
| Speech Dereverberation | spatialized WSJCAM0 | DeFT-AN | STOI | 0.981 | #1 of 1 | Archive leaderboard | report |
| Speech Enhancement | spatialized DNS challenge | DeFT-AN | PESQ | 3.01 | #1 of 1 | Archive leaderboard | report |
| Speech Enhancement | spatialized DNS challenge | DeFT-AN | SI-SDR | 9.9 | #1 of 1 | Archive leaderboard | report |
| Speech Enhancement | spatialized DNS challenge | DeFT-AN | STOI | 0.924 | #1 of 1 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections