Browse State-of-the-Art › Speaker Verification
Speaker Verification
200 papers with code · 12 benchmarks · 12 datasets archive 2025-07-28
Speaker verification is the verifying the identity of a person from characteristics of the voice.
( Image credit: Contrastive-Predictive-Coding-PyTorch )
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
12 leaderboard tables shown for this task, 12 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 12 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
12 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
3 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 200 papers with code (746 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 Oct 2017 30 repositories listed Syntology ran 4 of 27 samples · 23 unverified · 1 pointer-only (licence)In this paper, we propose a new loss function called generalized end-to-end (GE2E) loss, which makes the training of speaker verification models more efficient than our previous tuple-based end-to-end (TE2E) loss…
-
29 Jul 2018 26 repositories listed Syntology ran 6 of 9 samples · 3 unverified · 6 pointer-only (licence)Rather than employing standard hand-crafted features, the latter CNNs learn low-level speech representations from waveforms, potentially allowing the network to better capture important narrow-band speaker…
-
12 Jun 2018 11 repositories listedClone a voice in 5 seconds to generate arbitrary speech in real-time
-
5 Jun 2020 6 repositories listedTo explore this issue, we proposed to employ Mockingjay, a self-supervised learning based model, to protect anti-spoofing models against adversarial attacks in the black-box scenario.
-
17 Apr 2019 5 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)In this study, we explore end-to-end deep neural networks that input raw waveforms to improve various aspects: front-end speaker embedding extraction including model architecture, pre-training scheme, additional…
-
26 May 2017 5 repositories listedIn our paper, we propose an adaptive feature learning by utilizing the 3D-CNNs for direct speaker model creation in which, for both development and enrollment phases, an identical number of spoken utterances per speaker…
-
14 Apr 2019 4 repositories listed Syntology ran 3 of 6 samples · 3 unverifiedASVspoof, now in its third edition, is a series of community-led challenges which promote the development of countermeasures to protect automatic speaker verification (ASV) from the threat of spoofing.
-
5 Apr 2019 4 repositories listed Syntology ran 4 of 4 samples · 0 unverified · 3 pointer-only (licence)This paper proposes a novel unsupervised autoregressive neural model for learning generic speech representations.
-
28 Oct 2017 4 repositories listedFor many years, i-vector based audio embedding techniques were the dominant approach for speaker verification and speaker diarization applications.
-
24 Feb 2022 3 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedThe performance of spoofing countermeasure systems depends fundamentally upon the use of sufficiently representative training data.
-
3 Apr 2021 3 repositories listed Syntology ran 7 of 18 samples · 11 unverifiedSpoofing countermeasure (CM) systems are critical in speaker verification; they aim to discern spoofing attacks from bona fide speech trials.
-
27 Oct 2020 3 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedHuman voices can be used to authenticate the identity of the speaker, but the automatic speaker verification (ASV) systems are vulnerable to voice spoofing attacks, such as impersonation, replay, text-to-speech, and…
-
7 May 2020 3 repositories listedSpeaker recognition systems based on Convolutional Neural Networks (CNNs) are often built with off-the-shelf backbones such as VGG-Net or ResNet.
-
17 Sep 2019 3 repositories listedIn this work we present Ludwig, a flexible, extensible and easy to use toolbox which allows users to train deep learning models and use them for obtaining predictions without writing code.
-
9 Jun 2019 3 repositories listedIn the end, a posteriori SNR weighted energy difference is applied to the extended pitch segments of the denoised speech signal for detecting voice activity.
-
22 Sep 2017 3 repositories listedWe present a factorized hierarchical variational autoencoder, which learns disentangled and interpretable representations from sequential data without supervision.
-
27 Sep 2015 3 repositories listedIn this paper we present a data-driven, integrated approach to speaker verification, which maps a test utterance and a few reference utterances directly to a single score for verification and jointly optimizes the…
-
17 Jun 2024 2 repositories listedSDPN assigns the representation of the augmented views of an utterance to the same prototypes as the representation of the original view, thereby enabling effective knowledge transfer between the views.
-
29 Mar 2024 2 repositories listedWith 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis.
-
2 Mar 2024 2 repositories listedThis study proposes novel signal analysis methods for replay speech detection in automatic speaker verification (ASV) systems.
-
30 Jan 2024 2 repositories listed Syntology ran 9 of 15 samples · 6 unverified · 2 pointer-only (licence)First, we provide an open-source platform for researchers in the speaker recognition community to effortlessly build models.
-
13 Oct 2023 2 repositories listed Syntology ran 8 of 9 samples · 1 unverifiedIn this paper, we focus on a general yet important learning problem, pairwise similarity learning (PSL).
-
5 Aug 2023 2 repositories listedTo mitigate this problem, we introduce a diversity regularization term to embeddings in SDPN.
-
22 May 2023 2 repositories listedThis paper proposes a novel architecture called Enhanced Res2Net (ERes2Net), which incorporates both local and global feature fusion techniques to improve the performance.
-
1 May 2023 2 repositories listedThis paper describes the Ubenwa CryCeleb dataset - a labeled collection of infant cries - and the accompanying CryCeleb 2023 task, which is a public speaker verification challenge based on cry sounds.
-
4 Nov 2022 2 repositories listedOur previous research on one-class learning has improved the generalization ability to unseen attacks by compacting the bona fide speech in the embedding space.
-
14 Sep 2022 2 repositories listedWith the rapid development of speech conversion and speech synthesis algorithms, automatic speaker verification (ASV) systems are vulnerable to spoofing attacks.
-
20 May 2022 2 repositories listedPaddleSpeech is an open-source all-in-one speech toolkit.
-
23 Mar 2022 2 repositories listedParticipants apply their developed anonymization systems, run evaluation scripts and submit objective evaluation results and anonymized speech data to the organizers.
-
18 Mar 2022 2 repositories listed Syntology ran 0 of 7 samples · 7 unverifiedRecently, speech representation learning has improved many speech-related tasks such as speech recognition, speech classification, and speech-to-text translation.
Syntology lines on 11 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections