Datasets › VoxCeleb2
VoxCeleb2
VoxCeleb2 is a large scale speaker recognition dataset obtained automatically from open-source media. VoxCeleb2 consists of over a million utterances from over 6k speakers. Since the dataset is collected ‘in the wild’, the speech segments are corrupted with real world noise including laughter, cross-talk, channel effects, music and other sounds. The dataset is also multilingual, with speech from speakers of 145 different nationalities, covering a wide range of accents, ages, ethnicities and languages. The dataset is audio-visual, so is also useful for a number of other applications, for example – visual speech synthesis, speech separation, cross-modal transfer from face to voice or vice versa and training face recognition from video to complement existing face recognition datasets.
Source: VoxCeleb2: Deep Speaker Recognition Image Source: https://www.robots.ox.ac.uk/~vgg/data/voxceleb/
Benchmarks archive 2025-07-28
All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Speech Separation | VoxCeleb2 | IIANet SI-SNRi 14.0 | IIANet: An Intra- and Inter-Modality Attention Network... | JusperLee/IIANet | 5 | Compare |
| Talking Head Generation | VoxCeleb2 - 1-shot learning | Fast Bi-layer Avatars (medium size) CSIM 0.653 | Fast Bi-layer Neural Synthesis of One-Shot Realistic Head Avatars | saic-violet/bilayer-model | 5 | Compare |
| Talking Head Generation | VoxCeleb2 - 8-shot learning | CainGAN FID 24.9 | Pose Manipulation with Identity Preservation | — | 2 | Compare |
| Speaker Verification | VoxCeleb2 | ResNet-50 EER 100 | VoxCeleb2: Deep Speaker Recognition | a-nagrani/VGGVox +1 | 1 | Compare |
| Talking Head Generation | VoxCeleb2 - 32-shot learning | Few-shot Adversarial Model FID 30.6 | Few-Shot Adversarial Learning of Realistic Neural... | vincent-thevenin/Realistic-Neural-Talking-Head-Models +5 | 1 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 564. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| RTFS-Net: Recurrent Time-Frequency Modelling for Efficient Audio-Visual Speech Separation | 1 | 3 | 29 Sep 2023 | ran 0 of 1 samples (1 unverified) |
| IIANet: An Intra- and Inter-Modality Attention Network for Audio-Visual Speech Separation | 1 | 1 | 16 Aug 2023 | ran 1 of 1 samples (0 unverified) |
| An Audio-Visual Speech Separation Model Inspired by Cortico-Thalamo-Cortical Circuits | 2 | 1 | 21 Dec 2022 | ran 2 of 3 samples (1 unverified) |
| Fast Bi-layer Neural Synthesis of One-Shot Realistic Head Avatars | 1 | 3 | 24 Aug 2020 | not harvested |
| Pose Manipulation with Identity Preservation | 0 | 2 | 20 Apr 2020 | not harvested |
| Few-Shot Adversarial Learning of Realistic Neural Talking Head Models | 6 | 3 | 20 May 2019 | ran 2 of 2 samples (0 unverified; 2 pointer-only for licence) |
| VoxCeleb2: Deep Speaker Recognition | 2 | 1 | 14 Jun 2018 | ran 0 of 13 samples (13 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- VoxCeleb2 - 1-shot learning
- VoxCeleb2 - 8-shot learning
- VoxCeleb2 - 32-shot learning
- VoxCeleb2
4 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections