Datasets › VoxCeleb2

VoxCeleb2

Introduced by Joon Son Chung et al. in VoxCeleb2: Deep Speaker Recognition1 Jan 2018 archive 2025-07-28

VoxCeleb2 is a large scale speaker recognition dataset obtained automatically from open-source media. VoxCeleb2 consists of over a million utterances from over 6k speakers. Since the dataset is collected ‘in the wild’, the speech segments are corrupted with real world noise including laughter, cross-talk, channel effects, music and other sounds. The dataset is also multilingual, with speech from speakers of 145 different nationalities, covering a wide range of accents, ages, ethnicities and languages. The dataset is audio-visual, so is also useful for a number of other applications, for example – visual speech synthesis, speech separation, cross-modal transfer from face to voice or vice versa and training face recognition from video to complement existing face recognition datasets.

Source: VoxCeleb2: Deep Speaker Recognition Image Source: https://www.robots.ox.ac.uk/~vgg/data/voxceleb/

Benchmarks archive 2025-07-28

All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Speech Separation VoxCeleb2 IIANet SI-SNRi 14.0 IIANet: An Intra- and Inter-Modality Attention Network... JusperLee/IIANet 5 Compare
Talking Head Generation VoxCeleb2 - 1-shot learning Fast Bi-layer Avatars (medium size) CSIM 0.653 Fast Bi-layer Neural Synthesis of One-Shot Realistic Head Avatars saic-violet/bilayer-model 5 Compare
Talking Head Generation VoxCeleb2 - 8-shot learning CainGAN FID 24.9 Pose Manipulation with Identity Preservation — 2 Compare
Speaker Verification VoxCeleb2 ResNet-50 EER 100 VoxCeleb2: Deep Speaker Recognition a-nagrani/VGGVox +1 1 Compare
Talking Head Generation VoxCeleb2 - 32-shot learning Few-shot Adversarial Model FID 30.6 Few-Shot Adversarial Learning of Realistic Neural... vincent-thevenin/Realistic-Neural-Talking-Head-Models +5 1 Compare

Papers archive 2025-07-28

7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 564. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation 1 3 29 Sep 2023 ran 0 of 1 samples (1 unverified)
IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation 1 1 16 Aug 2023 ran 1 of 1 samples (0 unverified)
An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits 2 1 21 Dec 2022 ran 2 of 3 samples (1 unverified)
Fast Bi-layer Neural Synthesis of One-Shot Realistic Head Avatars 1 3 24 Aug 2020 not harvested
Pose Manipulation with Identity Preservation 0 2 20 Apr 2020 not harvested
Few-Shot Adversarial Learning of Realistic Neural Talking Head Models 6 3 20 May 2019 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
VoxCeleb2: Deep Speaker Recognition 2 1 14 Jun 2018 ran 0 of 13 samples (13 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • VoxCeleb2 - 1-shot learning
  • VoxCeleb2 - 8-shot learning
  • VoxCeleb2 - 32-shot learning
  • VoxCeleb2

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections