{"url":"/dataset/ava","name":"AVA","full_name":"Atomic Visual Actions","description_markdown":"**AVA** is a project that provides audiovisual annotations of video for improving our understanding of human activity. Each of the video clips has been exhaustively annotated by human annotators, and together they represent a rich variety of scenes, recording conditions, and expressions of human activity. There are annotations for:\r\n\r\n- Kinetics (AVA-Kinetics) - a crossover between AVA and Kinetics. In order to provide localized action labels on a wider variety of visual scenes, authors provide AVA action labels on videos from Kinetics-700, nearly doubling the number of total annotations, and increasing the number of unique videos by over 500x. \r\n- Actions (AvA Actions) - the AVA dataset densely annotates 80 atomic visual actions in 430 15-minute movie clips, where actions are localized in space and time, resulting in 1.62M action labels with multiple labels per human occurring frequently. \r\n- Spoken Activity (AVA ActiveSpeaker, AVA Speech). AVA ActiveSpeaker: associates speaking activity with a visible face, on the AVA v1.0 videos, resulting in 3.65 million frames labeled across ~39K face tracks. AVA Speech densely annotates audio-based speech activity in AVA v1.0 videos, and explicitly labels 3 background noise conditions, resulting in ~46K labeled segments spanning 45 hours of data.\r\nImage Source: [https://www.researchgate.net/profile/Paolo_Napoletano/publication/309327222/figure/fig1/AS:419620126248965@1477056642346/Sample-images-from-the-Aesthetic-Visual-Analysis-AVA-database-sorted-by-their-aesthetic.png](https://www.researchgate.net/profile/Paolo_Napoletano/publication/309327222/figure/fig1/AS:419620126248965@1477056642346/Sample-images-from-the-Aesthetic-Visual-Analysis-AVA-database-sorted-by-their-aesthetic.png)","description_withheld":null,"homepage":"http://research.google.com/ava/","introduced_date":"2018-01-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/ava-a-video-dataset-of-spatio-temporally","title":"AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions","first_author":"Chunhui Gu","url":null},"license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"}],"tasks":[{"name":"Action Recognition","url":"/task/action-recognition-in-videos","datasets_with_task":"/datasets/task/action-recognition-in-videos"},{"name":"Node Classification","url":"/task/node-classification","datasets_with_task":"/datasets/task/node-classification"},{"name":"Action Detection","url":"/task/action-detection","datasets_with_task":"/datasets/task/action-detection"},{"name":"Speech Enhancement","url":"/task/speech-enhancement","datasets_with_task":"/datasets/task/speech-enhancement"},{"name":"Action Recognition In Videos","url":"/task/action-recognition-in-videos-2","datasets_with_task":"/datasets/task/action-recognition-in-videos-2"},{"name":"Video Understanding","url":"/task/video-understanding","datasets_with_task":"/datasets/task/video-understanding"},{"name":"Speaker Diarization","url":"/task/speaker-diarization","datasets_with_task":"/datasets/task/speaker-diarization"},{"name":"Gaze Estimation","url":"/task/gaze-estimation","datasets_with_task":"/datasets/task/gaze-estimation"},{"name":"Self-Supervised Learning","url":"/task/self-supervised-learning","datasets_with_task":"/datasets/task/self-supervised-learning"},{"name":"Aesthetics Quality Assessment","url":"/task/aesthetics-quality-assessment","datasets_with_task":"/datasets/task/aesthetics-quality-assessment"},{"name":"Audio-Visual Active Speaker Detection","url":"/task/audio-visual-active-speaker-detection","datasets_with_task":"/datasets/task/audio-visual-active-speaker-detection"},{"name":"Spatio-Temporal Action Localization","url":"/task/spatio-temporal-action-localization","datasets_with_task":"/datasets/task/spatio-temporal-action-localization"},{"name":"Activity Detection","url":"/task/activity-detection","datasets_with_task":"/datasets/task/activity-detection"}],"languages":[],"variants":["AVA v2.1","AVA-ActiveSpeaker","AVA-LAEO","AVA-Speech","AVA v2.2","AVA-Kinetics"],"data_loaders":[{"repo":"https://github.com/open-mmlab/mmaction2","url":"https://github.com/open-mmlab/mmaction2/blob/master/tools/data/ava/README.md","frameworks":["pytorch"]}],"num_papers_in_archive":113,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/action-recognition-on-ava-v2-2","task":"Action Recognition","dataset_variant":"AVA v2.2","rows":38,"metrics":["mAP"],"first_row_in_archive_order":{"model":"LART (Hiera-H, K700 PT+FT)","paper":"/paper/on-the-benefits-of-3d-pose-and-tracking-for","metrics":{"mAP":"45.1"},"code_links":[{"title":"brjathu/LART","url":"https://github.com/brjathu/LART"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/action-recognition-in-videos-on-ava-v21","task":"Action Recognition","dataset_variant":"AVA v2.1","rows":15,"metrics":["mAP (Val)","GFlops","Params (M)"],"first_row_in_archive_order":{"model":"STAR/L","paper":"/paper/end-to-end-spatio-temporal-action","metrics":{"mAP (Val)":"41.7"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/aesthetics-quality-assessment-on-ava","task":"Aesthetics Quality Assessment","dataset_variant":"AVA","rows":9,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"MP_adam","paper":"/paper/attention-based-multi-patch-aggregation-for","metrics":{"Accuracy":"83.0%"},"code_links":[{"title":"Openning07/MPADA","url":"https://github.com/Openning07/MPADA"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/spatio-temporal-action-localization-on-ava","task":"Spatio-Temporal Action Localization","dataset_variant":"AVA-Kinetics","rows":7,"metrics":["val mAP","test mAP"],"first_row_in_archive_order":{"model":"VideoMAE V2-g","paper":"/paper/videomae-v2-scaling-video-masked-autoencoders","metrics":{"val mAP":"42.6"},"code_links":[{"title":"OpenGVLab/VideoMAEv2","url":"https://github.com/OpenGVLab/VideoMAEv2"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/node-classification-on-ava","task":"Node Classification","dataset_variant":"AVA","rows":4,"metrics":["mAP"],"first_row_in_archive_order":{"model":"ASDNet [ASDNet_ICCV2021]","paper":"/paper/learning-long-term-spatial-temporal-graphs","metrics":{"mAP":"93.5"},"code_links":[{"title":"sra2/spell","url":"https://github.com/sra2/spell"},{"title":"kylemin/SPELL","url":"https://github.com/kylemin/SPELL"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/action-recognition-in-videos-on-ava-v2-2","task":"Action Recognition In Videos","dataset_variant":"AVA v2.2","rows":2,"metrics":["mAP (Val)"],"first_row_in_archive_order":{"model":"YOWO+LFB*","paper":"/paper/you-only-watch-once-a-unified-cnn","metrics":{"mAP (Val)":"20.2"},"code_links":[{"title":"wei-tim/YOWO","url":"https://github.com/wei-tim/YOWO"},{"title":"zwtu/YOWO-Paddle","url":"https://github.com/zwtu/YOWO-Paddle"},{"title":"BoChenUIUC/YOWO","url":"https://github.com/BoChenUIUC/YOWO"},{"title":"nuschandra/Tennis-Stroke-Detection","url":"https://github.com/nuschandra/Tennis-Stroke-Detection"},{"title":"Stepphonwol/my_yowo","url":"https://github.com/Stepphonwol/my_yowo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/action-recognition-in-videos-on-ava-v2-1","task":"Action Recognition In Videos","dataset_variant":"AVA v2.1","rows":1,"metrics":["mAP (Val)"],"first_row_in_archive_order":{"model":"YOWO+LFB*","paper":"/paper/you-only-watch-once-a-unified-cnn","metrics":{"mAP (Val)":"19.2"},"code_links":[{"title":"wei-tim/YOWO","url":"https://github.com/wei-tim/YOWO"},{"title":"zwtu/YOWO-Paddle","url":"https://github.com/zwtu/YOWO-Paddle"},{"title":"BoChenUIUC/YOWO","url":"https://github.com/BoChenUIUC/YOWO"},{"title":"nuschandra/Tennis-Stroke-Detection","url":"https://github.com/nuschandra/Tennis-Stroke-Detection"},{"title":"Stepphonwol/my_yowo","url":"https://github.com/Stepphonwol/my_yowo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/asymmetric-masked-distillation-for-pre","title":"Asymmetric Masked Distillation for Pre-Training Small Foundation Models","date":"2023-11-06","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/hiera-a-hierarchical-vision-transformer","title":"Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles","date":"2023-06-01","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":0,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/end-to-end-spatio-temporal-action","title":"End-to-End Spatio-Temporal Action Localisation with Video Transformers","date":"2023-04-24","rows_on_this_dataset":3,"code_links":0,"syntology":null},{"paper":"/paper/on-the-benefits-of-3d-pose-and-tracking-for","title":"On the Benefits of 3D Pose and Tracking for Human Action Recognition","date":"2023-04-03","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/videomae-v2-scaling-video-masked-autoencoders","title":"VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking","date":"2023-03-29","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":2,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/unmasked-teacher-towards-training-efficient","title":"Unmasked Teacher: Towards Training-Efficient Video Foundation Models","date":"2023-03-28","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":3,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/masked-video-distillation-rethinking-masked","title":"Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning","date":"2022-12-08","rows_on_this_dataset":6,"code_links":4,"syntology":null},{"paper":"/paper/internvideo-general-video-foundation-models","title":"InternVideo: General Video Foundation Models via Generative and Discriminative Learning","date":"2022-12-06","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/holistic-interaction-transformer-network-for","title":"Holistic Interaction Transformer Network for Action Detection","date":"2022-10-23","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/learning-long-term-spatial-temporal-graphs","title":"Learning Long-Term Spatial-Temporal Graphs for Active Speaker Detection","date":"2022-07-15","rows_on_this_dataset":4,"code_links":2,"syntology":null},{"paper":"/paper/videomae-masked-autoencoders-are-data-1","title":"VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training","date":"2022-03-23","rows_on_this_dataset":8,"code_links":9,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":9,"samples_unverified":4,"pointer_only_for_licence":12,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/memvit-memory-augmented-multiscale-vision","title":"MeMViT: Memory-Augmented Multiscale Vision Transformer for Efficient Long-Term Video Recognition","date":"2022-01-20","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/masked-feature-prediction-for-self-supervised","title":"Masked Feature Prediction for Self-Supervised Visual Pre-Training","date":"2021-12-16","rows_on_this_dataset":1,"code_links":6,"syntology":null},{"paper":"/paper/improved-multiscale-vision-transformers-for","title":"MViTv2: Improved Multiscale Vision Transformers for Classification and Detection","date":"2021-12-02","rows_on_this_dataset":1,"code_links":9,"syntology":null},{"paper":"/paper/object-region-video-transformers-1","title":"Object-Region Video Transformers","date":"2021-10-13","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":1,"samples_unverified":6,"pointer_only_for_licence":7,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/towards-long-form-video-understanding-1","title":"Towards Long-Form Video Understanding","date":"2021-06-21","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":19,"samples_ran":5,"samples_unverified":14,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/relation-modeling-in-spatio-temporal-action","title":"Relation Modeling in Spatio-Temporal Action Localization","date":"2021-06-15","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/multiscale-vision-transformers","title":"Multiscale Vision Transformers","date":"2021-04-22","rows_on_this_dataset":6,"code_links":8,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":26,"samples_ran":13,"samples_unverified":13,"pointer_only_for_licence":5,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/pose-and-joint-aware-action-recognition","title":"Pose And Joint-Aware Action Recognition","date":"2020-10-16","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/actor-context-actor-relation-network-for","title":"Actor-Context-Actor Relation Network for Spatio-Temporal Action Localization","date":"2020-06-14","rows_on_this_dataset":4,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":27,"samples_ran":1,"samples_unverified":26,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/you-only-watch-once-a-unified-cnn","title":"You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization","date":"2019-11-15","rows_on_this_dataset":2,"code_links":5,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":3,"samples_unverified":9,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/effective-aesthetics-prediction-with-multi","title":"Effective Aesthetics Prediction with Multi-level Spatially Pooled Features","date":"2019-04-02","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/d3d-distilled-3d-networks-for-video-action","title":"D3D: Distilled 3D Networks for Video Action Recognition","date":"2018-12-19","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/long-term-feature-banks-for-detailed-video","title":"Long-Term Feature Banks for Detailed Video Understanding","date":"2018-12-12","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":0,"samples_unverified":8,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/slowfast-networks-for-video-recognition","title":"SlowFast Networks for Video Recognition","date":"2018-12-10","rows_on_this_dataset":8,"code_links":15,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":10,"samples_ran":0,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/video-action-transformer-network","title":"Video Action Transformer Network","date":"2018-12-06","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/attention-based-multi-patch-aggregation-for","title":"Attention-based Multi-Patch Aggregation for Image Aesthetic Assessment","date":"2018-10-22","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/actor-centric-relation-network","title":"Actor-Centric Relation Network","date":"2018-07-28","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/a-better-baseline-for-ava","title":"A Better Baseline for AVA","date":"2018-07-26","rows_on_this_dataset":2,"code_links":0,"syntology":null},{"paper":"/paper/nima-neural-image-assessment","title":"NIMA: Neural Image Assessment","date":"2017-09-15","rows_on_this_dataset":1,"code_links":12,"syntology":null},{"paper":"/paper/ava-a-video-dataset-of-spatio-temporally","title":"AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions","date":"2017-05-23","rows_on_this_dataset":1,"code_links":9,"syntology":null},{"paper":"/paper/a-lamp-adaptive-layout-aware-multi-patch-deep","title":"A-Lamp: Adaptive Layout-Aware Multi-Patch Deep Convolutional Neural Network for Photo Aesthetic Assessment","date":"2017-04-02","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/photo-aesthetics-ranking-network-with","title":"Photo Aesthetics Ranking Network with Attributes and Content Adaptation","date":"2016-06-06","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/composition-preserving-deep-photo-aesthetics","title":"Composition-Preserving Deep Photo Aesthetics Assessment","date":"2016-06-01","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/deep-aesthetic-quality-assessment-with","title":"Deep Aesthetic Quality Assessment with Semantic Information","date":"2016-04-18","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/deep-multi-patch-aggregation-network-for","title":"Deep Multi-Patch Aggregation Network for Image Style, Aesthetics, and Quality Estimation","date":"2015-12-01","rows_on_this_dataset":1,"code_links":0,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":12,"samples_harvested":145,"samples_ran":40,"samples_unverified":105,"pointer_only_for_licence":26,"papers_with_no_sample_that_ran":3,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}