Datasets › AVA-ActiveSpeaker

AVA-ActiveSpeaker

Introduced by Joseph Roth et al. in AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection archive 2025-07-28

Contains temporally labeled face tracks in video, where each face instance is labeled as speaking or not, and whether the speech is audible. This dataset contains about 3.65 million human labeled frames or about 38.5 hours of face tracks, and the corresponding audio.

Source: AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Audio-Visual Active Speaker Detection AVA-ActiveSpeaker LoCoNet+TalkNCE validation mean average precision 95.5% TalkNCE: Improving Active Speaker Detection with... kaistmm/TalkNCE 20 Compare

Papers archive 2025-07-28

18 shown of 18 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 22. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LASER: Lip Landmark Assisted Speaker Detection for Robustness 1 1 21 Jan 2025 not harvested
TalkNCE: Improving Active Speaker Detection with Talk-Aware Contrastive Learning 1 1 21 Sep 2023 not harvested
A Light Weight Model for Active Speaker Detection 1 1 8 Mar 2023 ran 5 of 6 samples (1 unverified)
LoCoNet: Long-Short Context Network for Active Speaker Detection 2 1 19 Jan 2023 not harvested
Audio-Visual Activity Guided Cross-Modal Identity Association for Active Speaker Detection 1 1 1 Dec 2022 not harvested
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection 2 2 15 Jul 2022 not harvested
UniCon+: ICTCAS-UCAS Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2022 0 1 22 Jun 2022 not harvested
End-to-End Active Speaker Detection 3 1 27 Mar 2022 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
Sub-word Level Lip Reading With Visual Attention 0 1 14 Oct 2021 not harvested
UniCon: Unified Context Network for Robust Active Speaker Detection 0 1 5 Aug 2021 not harvested
How to Design a Three-Stage Architecture for Audio-Visual Active Speaker Detection in the Wild 1 1 7 Jun 2021 not harvested
Active Speaker Detection as a Multi-Objective Optimization with Uncertainty-based Multimodal Fusion 0 1 7 Jun 2021 not harvested
NUS-HLT Report for ActivityNet Challenge 2021 AVA (Speaker) 1 1 1 Jun 2021 not harvested
ICTCAS-UCAS-TAL Submission to the AVA-ActiveSpeaker Task at ActivityNet Challenge 2021 0 1 1 Jun 2021 not harvested
MAAS: Multi-modal Assignation for Active Speaker Detection 1 2 11 Jan 2021 not harvested
Active Speakers in Context 1 1 20 May 2020 not harvested
Naver at ActivityNet Challenge 2019 -- Task B Active Speaker Detection (AVA) 0 1 25 Jun 2019 not harvested
Multi-Task Learning for Audio Visual Active Speaker Detection 0 1 1 Jun 2019 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • AVA-ActiveSpeaker

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections