{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/ava-a-video-dataset-of-spatio-temporally","title":"AVA: A Video Dataset of Spatio-temporally Localized Atomic Visual Actions","arxiv_id":"1705.08421","date":"2017-05-23","proceeding":"CVPR 2018 6","authors":["Chunhui Gu","Chen Sun","David A. Ross","Carl Vondrick","Caroline Pantofaru","Yeqing Li","Sudheendra Vijayanarasimhan","George Toderici","Susanna Ricco","Rahul Sukthankar","Cordelia Schmid","Jitendra Malik"],"abstract":"This paper introduces a video dataset of spatio-temporally localized Atomic\nVisual Actions (AVA). The AVA dataset densely annotates 80 atomic visual\nactions in 430 15-minute video clips, where actions are localized in space and\ntime, resulting in 1.58M action labels with multiple labels per person\noccurring frequently. The key characteristics of our dataset are: (1) the\ndefinition of atomic visual actions, rather than composite actions; (2) precise\nspatio-temporal annotations with possibly multiple annotations for each person;\n(3) exhaustive annotation of these atomic actions over 15-minute video clips;\n(4) people temporally linked across consecutive segments; and (5) using movies\nto gather a varied set of action representations. This departs from existing\ndatasets for spatio-temporal action recognition, which typically provide sparse\nannotations for composite actions in short video clips. We will release the\ndataset publicly.\n  AVA, with its realistic scene and action complexity, exposes the intrinsic\ndifficulty of action recognition. To benchmark this, we present a novel\napproach for action localization that builds upon the current state-of-the-art\nmethods, and demonstrates better performance on JHMDB and UCF101-24 categories.\nWhile setting a new state of the art on existing datasets, the overall results\non AVA are low at 15.6% mAP, underscoring the need for developing new\napproaches for video understanding.","url_abs":"http://arxiv.org/abs/1705.08421v4","url_pdf":"http://arxiv.org/pdf/1705.08421v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"ava-a-video-dataset-of-spatio-temporally","repo_url":"https://github.com/SmartPorridge/google-AVA-Dataset-downloader","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"unanswered"}},{"paper_slug":"ava-a-video-dataset-of-spatio-temporally","repo_url":"https://github.com/Whiffe/Custom-ava-dataset_Custom-Spatio-Temporally-Action-Video-Dataset","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"unanswered"}},{"paper_slug":"ava-a-video-dataset-of-spatio-temporally","repo_url":"https://github.com/tensorflow/models","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"tf","reach":null},{"paper_slug":"ava-a-video-dataset-of-spatio-temporally","repo_url":"https://github.com/2023-MindSpore-1/ms-code-17/tree/main/AVA_cifar","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":{"status":"ok"}},{"paper_slug":"ava-a-video-dataset-of-spatio-temporally","repo_url":"https://github.com/2023-MindSpore-4/Code8/tree/main/AVA_cifar","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"ava-a-video-dataset-of-spatio-temporally","repo_url":"https://github.com/Mind23-2/MindCode-11","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":{"status":"unanswered"}},{"paper_slug":"ava-a-video-dataset-of-spatio-temporally","repo_url":"https://github.com/MindSpore-paper-code-3/code6/tree/main/AVA_cifar","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null},{"paper_slug":"ava-a-video-dataset-of-spatio-temporally","repo_url":"https://github.com/open-mmlab/mmaction2","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"pytorch","reach":null},{"paper_slug":"ava-a-video-dataset-of-spatio-temporally","repo_url":"https://github.com/pwc-1/Paper-9/tree/main/3/AVA_cifar","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"mindspore","reach":null}],"tasks":[{"task_slug":null,"task_name":"Actin Detection"},{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"spatio-temporal-action-recognition","task_name":"Spatio-temporal Action Recognition"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[{"slug":"ava","name":"AVA","full_name":"Atomic Visual Actions"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-detection-on-j-hmdb","task":"Action Detection","dataset":"J-HMDB","model":"Faster-RCNN + two-stream I3D conv","rank_in_archive_order":7,"of":18,"metrics":{"Frame-mAP 0.5":"73.3","Video-mAP 0.5":"78.6"},"uses_additional_data":false},{"leaderboard":"/sota/action-detection-on-ucf101-24","task":"Action Detection","dataset":"UCF101-24","model":"Faster-RCNN + two-stream I3D conv","rank_in_archive_order":7,"of":19,"metrics":{"Frame-mAP 0.5":"76.3","Video-mAP 0.5":"59.9"},"uses_additional_data":false},{"leaderboard":"/sota/action-recognition-in-videos-on-ava-v21","task":"Action Recognition","dataset":"AVA v2.1","model":"S3D-G w/ ResNet RPN (Kinetics-400 pretraining(","rank_in_archive_order":13,"of":15,"metrics":{"mAP (Val)":"22.0"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1705.08421","atlas_url":"https://app.syntology.ai/?focus=1705.08421","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}