Papers › Cross-modal information fusion for voice spoofing detection

Cross-modal information fusion for voice spoofing detection

1 Feb 2023journal 2023 2archive 2025-07-28

Junxiao Xue, Hao Zhou, Huawei Song, Bin Wu, Lei Shi

In recent years, speaker verification systems have been used in many production scenarios. Unfortunately, they are still very vulnerable to different kinds of spoofing attacks, such as speech synthesis attacks, replay attacks, etc. Researchers have proposed many methods to defend against these attacks, but in the existing methods, researchers just focus on speech features. In recent studies, researchers have found that speech contains a large amount of face information. In fact, we can determine the speaker's gender, age, mouth shape, and other information by voice. These information can help us distinguish spoofing attacks. Inspired by this phenomenon, we propose a generalized framework named GACMNet. To cope with different attack scenarios, we instantiated two different models. Our framework is mainly divided into data pre-processing phase, feature extraction phase, feature fusion phase, and classification phase. Specifically, our framework consists of two branches. On the one hand, we extract face features in speech by a convolutional neural network. On the other hand, we use a densely connected network to extract speech features. For the more, we designed a global attention-based information fusion mechanism to distinguish the importance of each part of the features. Our solution was proven to be effective in two large scenarios. Compared to the existing methods, our model improves the tandem decision cost function (t-DCF) and equal error rate (EER) scores by 9% and 11% in the logical access scenario, respectively, our model improves the EER score by 10% in the physical access scenario.

PaperPDFCode

Code

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Automatic Speech RecognitionSpeaker VerificationSpeech SynthesisVoice Anti-spoofingfake voice detection

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Voice Anti-spoofing ASVspoof 2019 - LA LFCC&Face+SE-DenseNet+A-softmax EER 2.73 #4 of 8 Archive leaderboard report
Voice Anti-spoofing ASVspoof 2019 - LA LFCC&Face+SE-DenseNet+A-softmax min t-dcf 0.0713 #4 of 8 Archive leaderboard report
Voice Anti-spoofing ASVspoof 2019 - PA CQT&Face+SE-Res2Net+log-softmax EER 0.85 #1 of 1 Archive leaderboard report
Voice Anti-spoofing ASVspoof 2019 - PA CQT&Face+SE-Res2Net+log-softmax min t-dcf 0.0230 #1 of 1 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

1x1 ConvolutionAverage PoolingBatch NormalizationConcatenated Skip ConnectionConvolutionDense BlockDense ConnectionsDropoutGlobal Average PoolingKaiming InitializationMax PoolingReLURes2NetRes2Net BlockResidual ConnectionSoftmax

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections