{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/audio-deepfake-detection-with-self-supervised-1","title":"Audio Deepfake Detection with Self-Supervised XLS-R and SLS Classifier","arxiv_id":null,"date":"2024-10-28","proceeding":"ACM MM 2024 10","authors":["Qishan Zhang","Shuangbing Wen","Tao Hu"],"abstract":"Generative AI technologies, including text-to-speech (TTS) and\r\nvoice conversion (VC), frequently become indistinguishable from\r\ngenuine samples, posing challenges for individuals in discerning\r\nbetween real and synthetic content. This indistinguishability undermines trust in media, and the arbitrary cloning of personal voice\r\nsignals presents significant challenges to privacy and security. In\r\nthe field of deepfake audio detection, the majority of models achieving higher detection accuracy currently employ self-supervised\r\npre-trained models. However, with the ongoing development of\r\ndeepfake audio generation algorithms, maintaining high discrimination accuracy against new algorithms grows more challenging.\r\nTo enhance the sensitivity of deepfake audio features, we propose\r\na deepfake audio detection model that incorporates an SLS (Sensitive Layer Selection) module. Specifically, utilizing the pre-trained\r\nXLS-R enables our model to extract diverse audio features from its\r\nvarious layers, each providing distinct discriminative information.\r\nUtilizing the SLS classifier, our model captures sensitive contextual\r\ninformation across different layer levels of audio features, effectively\r\nemploying this information for fake audio detection. Experimental\r\nresults show that our method achieves state-of-the-art (SOTA) performance on both the ASVspoof 2021 DF and In-the-Wild datasets,\r\nwith a specific Equal Error Rate (EER) of 1.92% on the ASVspoof\r\n2021 DF dataset and 7.46% on the In-the-Wild dataset. Codes and\r\ndata can be found at https://github.com/QiShanZhang/SLSforADD.","url_abs":"https://openreview.net/pdf?id=acJMIXJg2u","url_pdf":"https://openreview.net/pdf?id=acJMIXJg2u","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"audio-deepfake-detection-with-self-supervised-1","repo_url":"https://github.com/qishanzhang/slsforadd","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"audio-deepfake-detection-with-self-supervised-1","repo_url":"https://github.com/QiShanZhang/SLSforASVspoof-2021-DF","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"audio-deepfake-detection","task_name":"Audio Deepfake Detection"},{"task_slug":"audio-generation","task_name":"Audio Generation"},{"task_slug":"deepfake-detection","task_name":"DeepFake Detection"},{"task_slug":"face-swapping","task_name":"Face Swapping"},{"task_slug":"text-to-speech","task_name":"Text to Speech"},{"task_slug":"voice-conversion","task_name":"Voice Conversion"},{"task_slug":"text-to-speech-1","task_name":"text-to-speech"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/audio-deepfake-detection-on-asvspoof-2021","task":"Audio Deepfake Detection","dataset":"ASVspoof 2021","model":"XLSR+SLS","rank_in_archive_order":3,"of":8,"metrics":{"21DF EER":"1.96","21LA EER":"2.86"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}