Home › Datasets › task › Action Recognition
Action Recognition datasets
archive 2025-07-28
115 datasets carry the task tag "Action Recognition" (the task itself: Action Recognition), ordered by the archive's paper count. Page 2 of 3: 48 shown of 115. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Action Recognition datasets 49–96 of 115
A database with 2,000 videos captured by surveillance cameras in real-world scenes.
16 papers · 1 benchmark
We present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects.
14 papers · 2 benchmarks
HAA500 (Human-Centric Atomic Action Dataset)
HAA500 is a manually annotated human-centric atomic action dataset for action recognition on 500 classes with over 591k labeled frames.
14 papers · 1 benchmark
The Watch-n-Patch dataset was created with the focus on modeling human activities, comprising multiple actions in a completely unsupervised setting.
13 papers · 0 benchmarks
YUP++ (YUP++ Dynamic Scenes dataset)
A new and challenging video database of dynamic scenes that more than doubles the size of those previously available.
13 papers · 1 benchmark
EPIC-SOUNDS is a large scale dataset of audio annotations capturing temporal extents and class labels within the audio stream of the egocentric videos from EPIC-KITCHENS-100.
12 papers · 2 benchmarks
EgoExoLearn is a fascinating dataset designed to bridge the gap between egocentric and exocentric views of procedural activities.
12 papers · 3 benchmarks
RareAct is a video dataset of unusual actions, including actions like “blend phone”, “cut keyboard” and “microwave shoes”.
12 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
11 papers · 1 benchmark
The AI City Challenge, hosted at CVPR 2024, focuses on harnessing AI to enhance operational efficiency in physical settings such as retail and warehouse environments, and Intelligent Traffic Systems (ITS).
10 papers · 1 benchmark
TAPOS is a new dataset developed on sport videos with manual annotations of sub-actions, and conduct a study on temporal action parsing on top.
10 papers · 1 benchmark
Home Action Genome is a large-scale multi-view video database of indoor daily activities.
9 papers · 2 benchmarks
We describe the 2020 edition of the DeepMind Kinetics human action dataset, which replenishes and extends the Kinetics-700 dataset.
9 papers · 1 benchmark
PSI-AVA is a dataset designed for holistic surgical scene understanding.
9 papers · 0 benchmarks
This is a 3D action recognition dataset, also known as 3D Action Pairs dataset.
7 papers · 1 benchmark
The Privacy Annotated HMDB51 (PA-HMDB51) dataset is a video-based dataset for evaluating pirvacy protection in visual action recognition algorithms.
7 papers · 0 benchmarks
The TUM Kitchen dataset is an action recognition dataset that contains 20 video sequences captured by 4 cameras with overlapping views.
7 papers · 0 benchmarks
This dataset contains 70 (30 falls + 40 activities of daily living) sequences.
7 papers · 0 benchmarks
ARID is a dataset for action recognition in dark videos.
6 papers · 0 benchmarks
IndustReal (IndustReal Dataset of Egocentric Videos for Procedure Understanding)
IndustReal is an ego-centric, multi-modal dataset where 27 participants are challenged to perform assembly and maintenance procedures on a construction-toy car.
6 papers · 3 benchmarks
This dataset has the following citation: M.
5 papers · 1 benchmark
RoCoG-v2 (Robot Control Gestures) is a dataset intended to support the study of synthetic-to-real and ground-to-air video domain adaptation.
5 papers · 1 benchmark
Simitate is a hybrid benchmarking suite targeting the evaluation of approaches for imitation learning.
5 papers · 0 benchmarks
UESTC RGB-D Varying-view action database contains 40 categories of aerobic exercise.
5 papers · 1 benchmark
TinyVIRAT contains natural low-resolution activities.
4 papers · 0 benchmarks
Verse is a new dataset that augments existing multimodal datasets (COCO and TUHOI) with sense labels.
4 papers · 0 benchmarks
CHAIRS is a large-scale motion-captured f-AHOI dataset, consisting of 17.3 hours of versatile interactions between 46 participants and 81 articulated and rigid sittable objects.
3 papers · 0 benchmarks
A novel dataset that represents complex conversational interactions between two individuals via 3D pose.
3 papers · 0 benchmarks
The Composable activities dataset consists of 693 videos that contain activities in 16 classes performed by 14 actors.
3 papers · 0 benchmarks
DCASE2014 is an audio classification benchmark.
3 papers · 0 benchmarks
Fitness-AQA (Fitness Action Quality Assessment [ECCV 2022])
Largest, first-of-its-kind, in-the-wild, fine-grained workout/exercise posture analysis dataset, covering three different exercises: BackSquat, Barbell Row, and Overhead Press.
3 papers · 0 benchmarks
A dataset for benchmarking action recognition algorithms in natural environments, while making use of 3D information.
3 papers · 0 benchmarks
IfAct (Identifying Human Actions Visible in Online Vlogs)
We consider the task of identifying human actions visible in online videos.
3 papers · 0 benchmarks
Is one of the largest egocentric datasets in the object search task with eyetracking information available Source: Deep Future Gaze: Gaze Anticipation on Egocentric Videos Using Adversarial Networks
3 papers · 0 benchmarks
PETRAW (PEg TRAnsfer Workflow recognition by different modalities)
PETRAW data set was composed of 150 sequences of peg transfer training sessions.
3 papers · 6 benchmarks
UAV-GESTURE is a dataset for UAV control and gesture recognition.
3 papers · 0 benchmarks
UESTC-MMEA-CL (A multi-modal egocentric activity dataset for continual learning)
UESTC-MMEA-CL is a new multi-modal activity dataset for continual egocentric activity recognition, which is proposed to promote future studies on continual learning for first-person activity recognition in wearable applications.
3 papers · 0 benchmarks
AVMIT (Audiovisual Moments in Time)
Audiovisual Moments in Time (AVMIT) is a large-scale dataset of audiovisual action events.
2 papers · 0 benchmarks
CVB (Video Dataset of Cattle Visual Behaviors)
Existing image/video datasets for cattle behavior recognition are mostly small, lack well-defined labels, or are collected in unrealistic controlled environments.
2 papers · 0 benchmarks
Drone-Action (Drone-Action: An Outdoor Recorded Drone Video Dataset for Action Recognition)
Website: https://asankagp.github.io/droneaction/
2 papers · 1 benchmark
This is a subset of Kinetics-400, introduced in Look, Listen and Learn by Relja Arandjelovic and Andrew Zisserman.
2 papers · 0 benchmarks
LoTE-Animal (LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior Understanding)
Understanding and analyzing animal behavior is increasingly essential to protect endangered animal species.
2 papers · 1 benchmark
MCAD (Multi-Camera Action Dataset)
Designed to evaluate the open view classification problem under the surveillance environment.
2 papers · 0 benchmarks
MPHOI-72 (Multi-person Human-object Interaction Dataset 72)
MPHOI-72 is a multi-person human-object interaction dataset that can be used for a wide variety of HOI/activity recognition and pose estimation/object tracking tasks.
2 papers · 0 benchmarks
MetaVD is a Meta Video Dataset for enhancing human action recognition datasets.
2 papers · 0 benchmarks
RISE is a large-scale video dataset for Recognizing Industrial Smoke Emissions.
2 papers · 0 benchmarks
A curated and 3-D pose-annotated subset of RGB videos sourced from Kinetics-700, a large-scale action dataset.
2 papers · 1 benchmark
A dataset derived from the recently introduced Mimetics dataset.
2 papers · 2 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.