Datasets › VoxCeleb1

VoxCeleb1

Introduced in VoxCeleb: a large-scale speaker identification dataset26 Jun 2017 archive 2025-07-28

VoxCeleb1 is an audio dataset containing over 100,000 utterances for 1,251 celebrities, extracted from videos uploaded to YouTube.

Benchmarks archive 2025-07-28

All 10 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Speaker Verification VoxCeleb SimAM-ResNet100 EER 0.20 VoxBlink2: A 100K+ Speaker Recognition Corpus and the... wenet-e2e/wespeaker 21 Compare
Speaker Verification VoxCeleb1 ReDimNet-B6-SF2-LM-ASNorm (15.0M) EER 0.37 Reshape Dimensions Network for Speaker Recognition IDRnD/ReDimNet 16 Compare
Speaker Identification VoxCeleb1 MSM-MAE Top-1 (%) 96.6 Masked Modeling Duo: Towards a Universal Audio... nttcslab/m2d +1 12 Compare
Few-Shot Audio Classification VoxCeleb1 Meta-Curvature (CRNN) Top-1 Accuracy(5-Way-1-Shot) 63.85 +- 0.44 MetaAudio: A Few-Shot Audio Classification Benchmark cheggan/metaaudio-a-few-shot-audio-classification-benchmark 10 Compare
Speaker Recognition VoxCeleb1 WavLM+ECAPA-TDNN EER 0.39 ESPnet-SPK: full pipeline speaker embedding toolkit with... espnet/espnet +1 2 Compare
Talking Head Generation VoxCeleb1 - 1-shot learning Few-shot Adversarial Model FID 43.0 Few-Shot Adversarial Learning of Realistic Neural... vincent-thevenin/Realistic-Neural-Talking-Head-Models +5 2 Compare
Talking Head Generation VoxCeleb1 - 32-shot learning Few-shot Adversarial Model FID 29.5 Few-Shot Adversarial Learning of Realistic Neural... vincent-thevenin/Realistic-Neural-Talking-Head-Models +5 2 Compare
Talking Head Generation VoxCeleb1 - 8-shot learning Few-shot Adversarial Model FID 38.0 Few-Shot Adversarial Learning of Realistic Neural... vincent-thevenin/Realistic-Neural-Talking-Head-Models +5 2 Compare
Video Reconstruction VoxCeleb Siarohin et al. AED 0.133 Motion Representations for Articulated Animation AliaksandrSiarohin/first-order-model +1 2 Compare
VoxCeleb1 M2D/0.7 Acc 96.3 Masked Modeling Duo: Towards a Universal Audio... nttcslab/m2d +1 1 Compare

Papers archive 2025-07-28

21 shown of 21 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 680. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Reshape Dimensions Network for Speaker Recognition 1 28 25 Jul 2024 not harvested
VoxBlink2: A 100K+ Speaker Recognition Corpus and the Open-Set Speaker-Identification Benchmark 1 1 16 Jul 2024 not harvested
SSAMBA: Self-Supervised Audio Representation Learning with Mamba State Space Model 1 1 20 May 2024 ran 5 of 8 samples (3 unverified)
Masked Modeling Duo: Towards a Universal Audio Pre-training Framework 2 4 9 Apr 2024 ran 7 of 7 samples (0 unverified; 7 pointer-only for licence)
ESPnet-SPK: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models 2 2 30 Jan 2024 ran 9 of 15 samples (6 unverified; 2 pointer-only for licence)
MT-SLVR: Multi-Task Self-Supervised Learning for Transformation In(Variant) Representations 1 3 29 May 2023 not harvested
Masked Modeling Duo: Learning Representations by Encouraging Both Networks to Model the Input 1 1 26 Oct 2022 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Masked Autoencoders that Listen 4 2 13 Jul 2022 ran 17 of 32 samples (15 unverified; 10 pointer-only for licence)
ATST: Audio Representation Learning with Teacher-Student Transformer 4 1 26 Apr 2022 not harvested
MetaAudio: A Few-Shot Audio Classification Benchmark 1 7 5 Apr 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding 0 1 4 Nov 2021 not harvested
SSAST: Self-Supervised Audio Spectrogram Transformer 3 2 19 Oct 2021 ran 11 of 16 samples (5 unverified)
TitaNet: Neural Model for speaker representation with 1D Depth-wise separable convolutions and global context 2 3 8 Oct 2021 ran 0 of 4 samples (4 unverified)
Multi-task Voice Activated Framework using Self-supervised Learning 0 1 3 Oct 2021 not harvested
Fine-tuning wav2vec2 for speaker recognition 4 1 30 Sep 2021 not harvested
SpeechNAS: Towards Better Trade-off between Latency and Accuracy for Large-Scale Speaker Verification 1 2 18 Sep 2021 not harvested
Motion Representations for Articulated Animation 2 2 22 Apr 2021 ran 2 of 6 samples (4 unverified; 5 pointer-only for licence)
Contrastive Learning of General-Purpose Audio Representations 2 1 21 Oct 2020 not harvested
AutoSpeech: Neural Architecture Search for Speaker Recognition 3 1 7 May 2020 not harvested
Few-Shot Adversarial Learning of Realistic Neural Talking Head Models 6 3 20 May 2019 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
X2Face: A network for controlling face generation using images, audio, and pose codes 0 3 1 Sep 2018 not harvested

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • VoxCeleb
  • VoxCeleb1
  • VoxCeleb1 - 1-shot learning
  • VoxCeleb1 - 8-shot learning
  • VoxCeleb1 - 32-shot learning

5 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections