Browse State-of-the-Art › Audio Deepfake Detection
Audio Deepfake Detection
35 papers with code · 2 benchmarks · 3 datasets archive 2025-07-28
Nowadays, deepfake is now generically used by the media or people to refer to any audio or video in which important attributes have been either digitally altered or swapped, with the help of artificial intelligence (AI). Audio deepfake detection is a task that aims to distinguish genuine utterances from fake ones via machine learning techniques.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 2 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| ASVspoof 2021 (8 rows) | XLSR-Mamba | XLSR-Mamba: A Dual-Column Bidirectional State Space Model for... | code | — | Compare |
| FakeOrReal (1 row) | rawnet_lite.pt | End-to-end Audio Deepfake Detection from RAW Waveforms: a... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 35 papers with code (74 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
24 Feb 2022 3 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedThe performance of spoofing countermeasure systems depends fundamentally upon the use of sufficiently representative training data.
-
28 Oct 2024 2 repositories listedTo enhance the sensitivity of deepfake audio features, we propose a deepfake audio detection model that incorporates an SLS (Sensitive Layer Selection) module.
-
26 Aug 2024 2 repositories listed Syntology ran 8 of 9 samples · 1 unverified · 9 pointer-only (licence)Existing research and datasets in fake song detection only focus on singing voice deepfake detection (SVDD), where the vocals are AI-generated but the instrumental music is sourced from real songs.
-
14 Aug 2024 2 repositories listedAudio deepfake detection has become a pivotal task over the last couple of years, as many recent speech synthesis and voice cloning systems generate highly realistic speech samples, thus enabling their use in malicious…
-
4 Nov 2021 2 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedDeep generative modeling has the potential to cause significant harm to society.
-
4 Oct 2021 2 repositories listed Syntology ran 11 of 12 samples · 1 unverifiedArtefacts that differentiate spoofed from bona-fide utterances can reside in spectral or temporal domains.
-
29 May 2025 1 repository listedRecent advancements in Text-to-Speech (TTS) models, particularly in voice cloning, have intensified the demand for adaptable and efficient deepfake detection methods.
-
29 Apr 2025 1 repository listedAudio deepfakes represent a growing threat to digital security and trust, leveraging advanced generative models to produce synthetic speech that closely mimics real human voices.
-
9 Apr 2025 1 repository listedThe rapid advancement of audio generation technologies has escalated the risks of malicious deepfake audio across speech, sound, singing voice, and music, threatening multimedia security and trust.
-
21 Mar 2025 1 repository listedDeepfakes have become a universal and rapidly intensifying concern of generative AI across various media types such as images, audio, and videos.
-
5 Feb 2025 1 repository listedThis paper conducts a comprehensive layer-wise analysis of self-supervised learning (SSL) models for audio deepfake detection across diverse contexts, including multilingual datasets (English, Chinese, Spanish),…
-
11 Jan 2025 1 repository listedCurrent research in audio deepfake detection is gradually transitioning from binary classification to multi-class tasks, referred as audio deepfake source tracing task.
-
16 Dec 2024 1 repository listedTo address this issue, we propose a continual learning method named Region-Based Optimization (RegO) for audio deepfake detection.
-
15 Nov 2024 1 repository listedTransformers and their variants have achieved great success in speech processing.
-
13 Oct 2024 1 repository listedWe study test-time domain adaptation for audio deepfake detection (ADD), addressing three challenges: (i) source-target domain gaps, (ii) limited target dataset size, and (iii) high computational costs.
-
Where are we in audio deepfake detection? A systematic analysis over generative and detection models6 Oct 2024 1 repository listedThrough extensive experiments, (1) we reveal the limitations of existing detection methods and demonstrate that foundation models exhibit stronger generalization capabilities, likely due to their model size and the…
-
14 Sep 2024 1 repository listedTo overcome the challenge of identifying diverse deepfake audio without semantic clues, we enhance our deepfake detector with real-world codec augmentation.
-
20 Aug 2024 1 repository listedCurrently, Audio Language Models (ALMs) are rapidly advancing due to the developments in large language models and audio neural codecs.
-
25 Jun 2024 1 repository listedRecent synthetic speech detectors leveraging the Transformer model have superior performance compared to the convolutional neural network counterparts.
-
24 Jun 2024 1 repository listedAs speech synthesis systems continue to make remarkable advances in recent years, the importance of robust deepfake detection systems that perform well in unseen systems has grown.
-
5 Jun 2024 1 repository listedFor effective OOD detection, we first explore current post-hoc OOD methods and propose NSD, a novel OOD approach in identifying novel deepfake algorithms through the similarity consideration of both feature and logits…
-
8 May 2024 1 repository listedWith the proliferation of Audio Language Model (ALM) based deepfake audio, there is an urgent need for generalized detection methods.
-
24 Apr 2024 1 repository listedThe detection models exhibited vulnerabilities, with FAR rising to 36.
-
31 Mar 2024 1 repository listedTo validate our hypothesis, we extract representations from state-of-the-art (SOTA) PTMs including monolingual, multilingual as well as PTMs trained for speaker and emotion recognition, and evaluated them on ASVSpoof…
-
21 Mar 2024 1 repository listedIn contrast to existing methods that fine-tune SSL models and employ additional deep neural networks for downstream tasks, we exploit classical machine learning algorithms such as logistic regression and shallow neural…
-
15 Dec 2023 1 repository listedThe rapid evolution of speech synthesis and voice conversion has raised substantial concerns due to the potential misuse of such technology, prompting a pressing need for effective audio deepfake detection mechanisms.
-
15 Sep 2023 1 repository listedAudio deepfake detection (ADD) is the task of detecting spoofing attacks generated by text-to-speech or voice conversion systems.
-
5 Sep 2023 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this paper, we initially construct a Chinese Fake Song Detection (FSD) dataset to investigate the field of song deepfake detection.
-
25 May 2023 1 repository listedIn this paper, we propose a novel ADD model, termed as M2S-ADD, that attempts to discover audio authenticity cues during the mono-to-stereo conversion process.
-
5 May 2023 1 repository listedVoice phishing (vishing) is increasingly popular due to the development of speech synthesis technology.
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections