Papers › TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion

TDFNet: An Efficient Audio-Visual Speech Separation Model with Top-down Fusion

25 Jan 2024arXiv:2401.14185archive 2025-07-28

Samuel Pegg, Kai Li, Xiaolin Hu

Audio-visual speech separation has gained significant traction in recent years due to its potential applications in various fields such as speech recognition, diarization, scene analysis and assistive technologies. Designing a lightweight audio-visual speech separation network is important for low-latency applications, but existing methods often require higher computational costs and more parameters to achieve better separation performance. In this paper, we present an audio-visual speech separation model called Top-Down-Fusion Net (TDFNet), a state-of-the-art (SOTA) model for audio-visual speech separation, which builds upon the architecture of TDANet, an audio-only speech separation method. TDANet serves as the architectural foundation for the auditory and visual networks within TDFNet, offering an efficient model with fewer parameters. On the LRS2-2Mix dataset, TDFNet achieves a performance increase of up to 10\% across all performance metrics compared with the previous SOTA method CTCNet. Remarkably, these results are achieved using fewer parameters and only 28\% of the multiply-accumulate operations (MACs) of CTCNet. In essence, our method presents a highly effective and efficient solution to the challenges of speech separation within the audio-visual domain, making significant strides in harnessing visual information optimally.

PaperPDFCode

Code

spkgyk/TDFNet officialpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Speech RecognitionSpeech Separationspeech-recognition

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Separation LRS2 TDFNet-large PESQ 3.21 #2 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet-large SDRi 15.9 #2 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet-large SI-SNRi 15.8 #2 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet-large STOI 0.949 #2 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet (MHSA + Shared) PESQ 3.16 #3 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet (MHSA + Shared) SDRi 15.2 #3 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet (MHSA + Shared) SI-SNRi 15.0 #3 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet (MHSA + Shared) STOI 0.938 #3 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet-small PESQ 3.10 #8 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet-small SDRi 13.7 #8 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet-small SI-SNRi 13.6 #8 of 8 Archive leaderboard report
Speech Separation LRS2 TDFNet-small STOI 0.931 #8 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections