{"url":"/dataset/asvspoof-2021","name":"ASVspoof 2021","full_name":"ASVspoof 2021 Dataset","description_markdown":"Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab conditions towards to those encountered in the wild. ASVspoof, the spoofing and deepfake detection initiative and challenge series, has followed the same trend. This article provides a summary of the ASVspoof 2021 challenge and the results of 54 participating teams that submitted to the evaluation phase. For the logical access (LA) task, results indicate that countermeasures are robust to newly introduced encoding and transmission effects. Results for the physical access (PA) task indicate the potential to detect replay attacks in real, as opposed to simulated physical spaces, but a lack of robustness to variations between simulated and real acoustic environments. The Deepfake (DF) task, new to the 2021 edition, targets solutions to the detection of manipulated, compressed speech data posted online. While detection solutions offer some resilience to compression effects, they lack generalization across different source datasets. In addition to a summary of the top-performing systems for each task, new analyses of influential data factors and results for hidden data subsets, the article includes a review of post-challenge results, an outline of the principal challenge limitations and a road-map for the future of ASVspoof.","description_withheld":null,"homepage":"https://www.asvspoof.org/index2021.html","introduced_date":null,"introduced_date_note":null,"introduced_by":null,"license":null,"modalities":[],"tasks":[{"name":"Audio Deepfake Detection","url":"/task/audio-deepfake-detection","datasets_with_task":"/datasets/task/audio-deepfake-detection"}],"languages":[],"variants":["ASVspoof 2021"],"data_loaders":[],"num_papers_in_archive":8,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/audio-deepfake-detection-on-asvspoof-2021","task":"Audio Deepfake Detection","dataset_variant":"ASVspoof 2021","rows":8,"metrics":["21LA EER","21DF EER"],"first_row_in_archive_order":{"model":"XLSR-Mamba","paper":"/paper/xlsr-mamba-a-dual-column-bidirectional-state","metrics":{"21DF EER":"1.88","21LA EER":"0.93"},"code_links":[{"title":"swagshaw/xlsr-mamba","url":"https://github.com/swagshaw/xlsr-mamba"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/xlsr-mamba-a-dual-column-bidirectional-state","title":"XLSR-Mamba: A Dual-Column Bidirectional State Space Model for Spoofing Attack Detection","date":"2024-11-15","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/audio-deepfake-detection-with-self-supervised-1","title":"Audio Deepfake Detection with Self-Supervised XLS-R and SLS Classifier","date":"2024-10-28","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/temporal-channel-modeling-in-multi-head-self","title":"Temporal-Channel Modeling in Multi-head Self-Attention for Synthetic Speech Detection","date":"2024-06-25","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/bts-e-audio-deepfake-detection-using","title":"Bts-e: Audio deepfake detection using breathing-talking-silence encoder","date":"2023-05-05","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/automatic-speaker-verification-spoofing-and","title":"Automatic speaker verification spoofing and deepfake detection using wav2vec 2.0 and data augmentation","date":"2022-02-24","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":1,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/aasist-audio-anti-spoofing-using-integrated","title":"AASIST: Audio Anti-Spoofing using Integrated Spectro-Temporal Graph Attention Networks","date":"2021-10-04","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":11,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/end-to-end-spectro-temporal-graph-attention","title":"End-to-End Spectro-Temporal Graph Attention Networks for Speaker Verification Anti-Spoofing and Speech Deepfake Detection","date":"2021-07-27","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/end-to-end-anti-spoofing-with-rawnet2","title":"End-to-end anti-spoofing with RawNet2","date":"2020-11-02","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":3,"samples_harvested":17,"samples_ran":15,"samples_unverified":2,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}