Browse State-of-the-Art › Speaker Recognition
Speaker Recognition
102 papers with code · 1 benchmark · 6 datasets archive 2025-07-28
Speaker Recognition is the process of identifying or confirming the identity of a person given his speech segments.
Source: Margin Matters: Towards More Discriminative Deep Neural Network Embeddings for Speaker Recognition
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
1 leaderboard table shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| VoxCeleb1 (2 rows) | WavLM+ECAPA-TDNN | ESPnet-SPK: full pipeline speaker embedding toolkit with... | code | Syntology ran 9 of 15 samples · 6 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 102 papers with code (435 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Jul 2018 26 repositories listed Syntology ran 6 of 9 samples · 3 unverified · 6 pointer-only (licence)Rather than employing standard hand-crafted features, the latter CNNs learn low-level speech representations from waveforms, potentially allowing the network to better capture important narrow-band speaker…
-
5 May 2017 15 repositories listed Syntology ran 7 of 9 samples · 2 unverified · 5 pointer-only (licence)We present Deep Speaker, a neural speaker embedding system that maps utterances to a hypersphere where speaker similarity is measured by cosine similarity.
-
26 Feb 2019 10 repositories listed Syntology ran 3 of 23 samples · 20 unverified · 1 pointer-only (licence)The objective of this paper is speaker recognition "in the wild"-where utterances may be of variable length and also contain irrelevant signals.
-
12 Jul 2020 7 repositories listed Syntology ran 5 of 14 samples · 9 unverifiedWe present a large-scale comparison of various self-supervised models.
-
Mockingjay: Unsupervised Speech Representation Learning with Deep Bidirectional Transformer Encoders25 Oct 2019 7 repositories listed Syntology ran 3 of 11 samples · 8 unverifiedWe present Mockingjay as a new speech representation learning approach, where bidirectional Transformer encoders are pre-trained on a large amount of unlabeled speech.
-
13 Oct 2020 6 repositories listedTo this end, in this work, we present a novel end-to-end 3D lip motion Network (3LMNet) by utilizing the sentence-level 3D lip motion (S3DLM) to recognize speakers in both the text-independent and text-dependent…
-
11 Oct 2018 5 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this paper, we present a novel system that separates the voice of a target speaker from multi-speaker signals, by making use of a reference signal from the target speaker.
-
30 Sep 2021 4 repositories listedThis paper explores applying the wav2vec2 framework to speaker recognition instead of speech recognition.
-
28 Mar 2022 3 repositories listedIn speaker recognition, where speech segments are mapped to embeddings on the unit hypersphere, two scoring backends are commonly used, namely cosine scoring or PLDA.
-
12 Nov 2021 3 repositories listedThis work provides a brief description of Human Language Technology (HLT) Laboratory, National University of Singapore (NUS) system submission for 2020 NIST conversational telephone speech (CTS) speaker recognition…
-
7 May 2020 3 repositories listedSpeaker recognition systems based on Convolutional Neural Networks (CNNs) are often built with off-the-shelf backbones such as VGG-Net or ResNet.
-
31 Mar 2020 3 repositories listedTo address this demand, we propose a portable model called Additive Margin MobileNet1D (AM-MobileNet1D) to Speaker Identification on mobile devices.
-
29 Mar 2024 2 repositories listedWith 3D-Speaker-Toolkit, we establish a new benchmark for multimodal speaker analysis.
-
30 Jan 2024 2 repositories listed Syntology ran 9 of 15 samples · 6 unverified · 2 pointer-only (licence)First, we provide an open-source platform for researchers in the speaker recognition community to effortlessly build models.
-
16 Feb 2023 2 repositories listedWe propose an objective for perceptual quality based on temporal acoustic parameters.
-
27 Oct 2022 2 repositories listedIt extends PSDA with the ability to model within and between-speaker variabilities in toroidal submanifolds of the hypersphere.
-
7 Jun 2022 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)According to the characteristic of SRSs, we present 22 diverse transformations and thoroughly evaluate them using 7 recent promising adversarial attacks (4 white-box and 3 black-box) on speaker recognition.
-
20 May 2022 2 repositories listedPaddleSpeech is an open-source all-in-one speech toolkit.
-
16 Mar 2022 2 repositories listedOur best model achieves an equal error rate of 0.
-
2 Nov 2020 2 repositories listedAnonymisation has the goal of manipulating speech signals in order to degrade the reliability of automatic approaches to speaker recognition, while preserving other aspects of speech, such as those relating to…
-
31 May 2020 2 repositories listedTime Delay Neural Network (TDNN) is a well-performing structure for DNN-based speaker recognition systems.
-
25 Feb 2020 2 repositories listedWe compare the three best architectures trained using our method to select the best one, which is the one with a shallow architecture.
-
31 Oct 2019 2 repositories listedThese datasets tend to deliver over optimistic performance and do not meet the request of research on speaker recognition in unconstrained conditions.
-
23 Oct 2019 2 repositories listed Syntology ran 1 of 8 samples · 7 unverifiedAlso, we validate the use of parameterized filterbanks and show that complex-valued representations and masks are beneficial in all conditions.
-
12 Aug 2019 2 repositories listedIn this paper, we propose "personal VAD", a system to detect the voice activity of a target speaker at the frame level.
-
13 Dec 2018 2 repositories listedDeep neural networks can learn complex and abstract representations, that are progressively obtained by combining simpler ones.
-
14 Jun 2018 2 repositories listed Syntology ran 0 of 13 samples · 13 unverifiedThe objective of this paper is speaker recognition under noisy and unconstrained conditions.
-
28 Oct 2017 2 repositories listedAttention-based models have recently shown great performance on a range of tasks, such as speech recognition, machine translation, and image captioning due to their ability to summarize relevant information that expands…
-
11 Jun 2025 1 repository listedSpeaker recognition systems are often limited to classification tasks and struggle to generate detailed speaker characteristics or provide context-rich descriptions.
-
23 May 2025 1 repository listedSpeaker anonymization seeks to conceal a speaker's identity while preserving the utility of their speech.
Syntology lines on 10 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections