Browse State-of-the-Art › Speaker Identification
Speaker Identification
74 papers with code · 4 benchmarks · 4 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
4 leaderboard tables shown for this task, 4 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| VoxCeleb1 (12 rows) | MSM-MAE | Masked Modeling Duo: Towards a Universal Audio Pre-training Framework | code | Syntology ran 7 of 7 samples · 0 unverified | Compare |
| EVI en-GB (1 row) | Fuzzy Retrieval | EVI: Multilingual Spoken Dialogue Tasks and Dataset for... | code | — | Compare |
| EVI fr-FR (1 row) | Fuzzy Retrieval | EVI: Multilingual Spoken Dialogue Tasks and Dataset for... | code | — | Compare |
| EVI pl-PL (1 row) | Fuzzy Retrieval | EVI: Multilingual Spoken Dialogue Tasks and Dataset for... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
4 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 74 papers with code (248 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
29 Jul 2018 26 repositories listed Syntology ran 6 of 9 samples · 3 unverified · 6 pointer-only (licence)Rather than employing standard hand-crafted features, the latter CNNs learn low-level speech representations from waveforms, potentially allowing the network to better capture important narrow-band speaker…
-
5 May 2017 15 repositories listed Syntology ran 7 of 9 samples · 2 unverified · 5 pointer-only (licence)We present Deep Speaker, a neural speaker embedding system that maps utterances to a hypersphere where speaker similarity is measured by cosine similarity.
-
14 Oct 2021 6 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)Motivated by the success of T5 (Text-To-Text Transfer Transformer) in pre-trained natural language processing models, we propose a unified-modal SpeechT5 framework that explores the encoder-decoder pre-training for…
-
12 Oct 2021 5 repositories listedWe integrate the proposed methods into the HuBERT framework.
-
13 Jul 2022 4 repositories listed Syntology ran 17 of 32 samples · 15 unverified · 10 pointer-only (licence)Following the Transformer encoder-decoder design in MAE, our Audio-MAE first encodes audio spectrogram patches with a high masking ratio, feeding only the non-masked tokens through encoder layers.
-
26 Apr 2022 4 repositories listedSelf-supervised learning (SSL) learns knowledge from a large amount of unlabeled data, and then transfers the knowledge to a specific problem with a limited number of labeled data.
-
18 May 2020 4 repositories listedWe use the representations with two downstream tasks, speaker identification, and phoneme classification.
-
19 Oct 2021 3 repositories listed Syntology ran 11 of 16 samples · 5 unverifiedHowever, pure Transformer models tend to require more training data compared to CNNs, and the success of the AST relies on supervised pretraining that requires a large amount of labeled data and a complex training…
-
7 May 2020 3 repositories listedSpeaker recognition systems based on Convolutional Neural Networks (CNNs) are often built with off-the-shelf backbones such as VGG-Net or ResNet.
-
31 Mar 2020 3 repositories listedTo address this demand, we propose a portable model called Additive Margin MobileNet1D (AM-MobileNet1D) to Speaker Identification on mobile devices.
-
24 Sep 2024 2 repositories listedThe comic domain is rapidly advancing with the development of single- and multi-page analysis and synthesis models.
-
9 Apr 2024 2 repositories listed Syntology ran 7 of 7 samples · 0 unverified · 7 pointer-only (licence)This study proposes Masked Modeling Duo (M2D), an improved masked prediction SSL, which learns by predicting representations of masked input signals that serve as training signals.
-
20 May 2022 2 repositories listedPaddleSpeech is an open-source all-in-one speech toolkit.
-
1 Apr 2021 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We propose using self-supervised discrete representations for the task of speech resynthesis.
-
17 Nov 2020 2 repositories listedSpeaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification.
-
21 Oct 2020 2 repositories listedWe introduce COLA, a self-supervised pre-training approach for learning a general-purpose representation of audio.
-
25 Feb 2020 2 repositories listedWe compare the three best architectures trained using our method to select the best one, which is the one with a shallow architecture.
-
23 Oct 2019 2 repositories listedLearning meaningful and general representations from unannotated speech that are applicable to a wide range of tasks remains challenging.
-
22 Oct 2019 2 repositories listedRecent breakthroughs in deep learning often rely on representation learning and knowledge transfer.
-
1 Dec 2018 2 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedMutual Information (MI) or similar measures of statistical dependence are promising tools for learning these representations in an unsupervised way.
-
11 Jun 2025 1 repository listedSpeaker recognition systems are often limited to classification tasks and struggle to generate detailed speaker characteristics or provide context-rich descriptions.
-
13 Mar 2025 1 repository listedSpeaker identification in multilingual settings presents unique challenges, particularly when conventional models are predominantly trained on English data.
-
23 Dec 2024 1 repository listedBased on this Friends-MMC dataset, we further study two fundamental MMC tasks: conversation speaker identification and conversation response prediction, both of which have the multi-party nature with the video or image…
-
22 Nov 2024 1 repository listedVoice recognition and speaker identification are vital for applications in security and personal assistants.
-
3 Oct 2024 1 repository listedNeural speech models build deeply entangled internal representations, which capture a variety of features (e.
-
7 Sep 2024 1 repository listedHowever, after carefully examining Gaokao's questions, we find the correct answers to many questions can be inferred from the conversation transcript alone, i.
-
13 Aug 2024 1 repository listedIn the fields of security systems, forensic investigations, and personalized services, the importance of speech as a fundamental human input outweighs text-based interactions.
-
Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models16 Jul 2024 1 repository listedWe introduce an approach to identifying speaker names in dialogue transcripts, a crucial task for enhancing content accessibility and searchability in digital media archives.
-
16 Jul 2024 1 repository listedIn this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild.
-
4 Jul 2024 1 repository listedWe introduce a novel benchmark, CoMix, designed to evaluate the multi-task capabilities of models in comic analysis.
Syntology lines on 8 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections