Datasets › AVA

AVA (Atomic Visual Actions)

Introduced by Chunhui Gu et al. in AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions1 Jan 2018 archive 2025-07-28

AVA is a project that provides audiovisual annotations of video for improving our understanding of human activity. Each of the video clips has been exhaustively annotated by human annotators, and together they represent a rich variety of scenes, recording conditions, and expressions of human activity. There are annotations for:

  • Kinetics (AVA-Kinetics) - a crossover between AVA and Kinetics. In order to provide localized action labels on a wider variety of visual scenes, authors provide AVA action labels on videos from Kinetics-700, nearly doubling the number of total annotations, and increasing the number of unique videos by over 500x.
  • Actions (AvA Actions) - the AVA dataset densely annotates 80 atomic visual actions in 430 15-minute movie clips, where actions are localized in space and time, resulting in 1.62M action labels with multiple labels per human occurring frequently.
  • Spoken Activity (AVA ActiveSpeaker, AVA Speech). AVA ActiveSpeaker: associates speaking activity with a visible face, on the AVA v1.0 videos, resulting in 3.65 million frames labeled across ~39K face tracks. AVA Speech densely annotates audio-based speech activity in AVA v1.0 videos, and explicitly labels 3 background noise conditions, resulting in ~46K labeled segments spanning 45 hours of data. Image Source: https://www.researchgate.net/profile/Paolo_Napoletano/publication/309327222/figure/fig1/AS:419620126248965@1477056642346/Sample-images-from-the-Aesthetic-Visual-Analysis-AVA-database-sorted-by-their-aesthetic.png

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 36 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 113. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Asymmetric Masked Distillation for Pre-Training Small Foundation Models 0 1 6 Nov 2023 not harvested
Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles 4 1 1 Jun 2023 ran 0 of 6 samples (6 unverified)
End-to-End Spatio-Temporal Action Localisation with Video Transformers 0 3 24 Apr 2023 not harvested
On the Benefits of 3D Pose and Tracking for Human Action Recognition 1 1 3 Apr 2023 not harvested
VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking 1 3 29 Mar 2023 ran 2 of 6 samples (4 unverified)
Unmasked Teacher: Towards Training-Efficient Video Foundation Models 1 1 28 Mar 2023 ran 3 of 8 samples (5 unverified)
Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning 4 6 8 Dec 2022 not harvested
InternVideo: General Video Foundation Models via Generative and Discriminative Learning 2 2 6 Dec 2022 ran 3 of 3 samples (0 unverified)
Holistic Interaction Transformer Network for Action Detection 1 1 23 Oct 2022 not harvested
Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection 2 4 15 Jul 2022 not harvested
VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training 9 8 23 Mar 2022 ran 9 of 13 samples (4 unverified; 12 pointer-only for licence)
MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition 1 1 20 Jan 2022 not harvested
Masked Feature Prediction for Self-Supervised Visual Pre-Training 6 1 16 Dec 2021 not harvested
MViTv2: Improved Multiscale Vision Transformers for Classification and Detection 9 1 2 Dec 2021 not harvested
Object-Region Video Transformers 1 1 13 Oct 2021 ran 1 of 7 samples (6 unverified; 7 pointer-only for licence)
Towards Long-Form Video Understanding 2 1 21 Jun 2021 ran 5 of 19 samples (14 unverified)
Relation Modeling in Spatio-Temporal Action Localization 0 2 15 Jun 2021 not harvested
Multiscale Vision Transformers 8 6 22 Apr 2021 ran 13 of 26 samples (13 unverified; 5 pointer-only for licence)
Pose And Joint-Aware Action Recognition 1 1 16 Oct 2020 not harvested
Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization 3 4 14 Jun 2020 ran 1 of 27 samples (26 unverified)
You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization 5 2 15 Nov 2019 ran 3 of 12 samples (9 unverified; 2 pointer-only for licence)
Effective Aesthetics Prediction with Multi-level Spatially Pooled Features 1 1 2 Apr 2019 not harvested
D3D: Distilled 3D Networks for Video Action Recognition 1 1 19 Dec 2018 not harvested
Long-Term Feature Banks for Detailed Video Understanding 4 1 12 Dec 2018 ran 0 of 8 samples (8 unverified)
SlowFast Networks for Video Recognition 15 8 10 Dec 2018 ran 0 of 10 samples (10 unverified)
Video Action Transformer Network 0 2 6 Dec 2018 not harvested
Attention-based Multi-Patch Aggregation for Image Aesthetic Assessment 1 1 22 Oct 2018 not harvested
Actor-Centric Relation Network 1 1 28 Jul 2018 not harvested
A Better Baseline for AVA 0 2 26 Jul 2018 not harvested
NIMA: Neural Image Assessment 12 1 15 Sep 2017 not harvested

The full list of 36 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • AVA v2.1
  • AVA-ActiveSpeaker
  • AVA-LAEO
  • AVA-Speech
  • AVA v2.2
  • AVA-Kinetics

6 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections