Papers › PrimeK-Net: Multi-scale Spectral Learning via Group Prime-Kernel Convolutional Neural...

PrimeK-Net: Multi-scale Spectral Learning via Group Prime-Kernel Convolutional Neural Networks for Single Channel Speech Enhancement

27 Feb 2025arXiv:2502.19906archive 2025-07-28

Zizhen Lin, Junyu Wang, Ruili Li, Fei Shen, Xi Xuan

Single-channel speech enhancement is a challenging ill-posed problem focused on estimating clean speech from degraded signals. Existing studies have demonstrated the competitive performance of combining convolutional neural networks (CNNs) with Transformers in speech enhancement tasks. However, existing frameworks have not sufficiently addressed computational efficiency and have overlooked the natural multi-scale distribution of the spectrum. Additionally, the potential of CNNs in speech enhancement has yet to be fully realized. To address these issues, this study proposes a Deep Separable Dilated Dense Block (DSDDB) and a Group Prime Kernel Feedforward Channel Attention (GPFCA) module. Specifically, the DSDDB introduces higher parameter and computational efficiency to the Encoder/Decoder of existing frameworks. The GPFCA module replaces the position of the Conformer, extracting deep temporal and frequency features of the spectrum with linear complexity. The GPFCA leverages the proposed Group Prime Kernel Feedforward Network (GPFN) to integrate multi-granularity long-range, medium-range, and short-range receptive fields, while utilizing the properties of prime numbers to avoid periodic overlap effects. Experimental results demonstrate that PrimeK-Net, proposed in this study, achieves state-of-the-art (SOTA) performance on the VoiceBank+Demand dataset, reaching a PESQ score of 3.61 with only 1.41M parameters.

PaperPDFCode

Code

huaidanquede/PrimeK-Net officialmentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Computational EfficiencySpeech Enhancement

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Speech Enhancement VoiceBank + DEMAND PrimeK-Net CBAK 3.98 #7 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND PrimeK-Net COVL 4.35 #7 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND PrimeK-Net CSIG 4.81 #7 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND PrimeK-Net PESQ (wb) 3.61 #7 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND PrimeK-Net Para. (M) 1.41 #7 of 42 Archive leaderboard report
Speech Enhancement VoiceBank + DEMAND PrimeK-Net STOI 96 #7 of 42 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

AttentionBatch NormalizationConcatenated Skip ConnectionConvolutionDense BlockDense ConnectionsFeedforward NetworkReLUSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections