Home › Datasets › task › Action Recognition

Action Recognition datasets

archive 2025-07-28

115 datasets carry the task tag "Action Recognition" (the task itself: Action Recognition), ordered by the archive's paper count. Page 1 of 3: 48 shown of 115. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 50 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Action Recognition datasets 1–48 of 115

UCF101 (UCF101 Human Actions dataset)
UCF101 dataset is an extension of UCF50 and consists of 13,320 video clips, which are classified into 101 categories.
1,863 papers · 23 benchmarks
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
The HMDB51 dataset is a large collection of realistic videos from various sources, including movies and web videos.
839 papers · 10 benchmarks
The ActivityNet dataset contains 200 different types of activities and a total of 849 hours of videos collected from YouTube.
807 papers · 17 benchmarks
The Charades dataset is composed of 9,848 videos of daily indoors activities with an average length of 30 seconds, involving interactions with 46 objects classes in 15 types of indoor scenes and containing a vocabulary of 30 verbs leading…
428 papers · 6 benchmarks
The THUMOS14 (THUMOS 2014) dataset is a large-scale video dataset that includes 1,010 videos for validation and 1,574 videos for testing from 20 classes.
318 papers · 18 benchmarks
The 20BN-SOMETHING-SOMETHING V2 dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
290 papers · 7 benchmarks
HowTo100M is a large-scale dataset of narrated videos with an emphasis on instructional videos where content creators teach complex tasks with an explicit intention of explaining the visual content on screen.
286 papers · 1 benchmark
KTH (KTH Action dataset)
The efforts to create a non-trivial and publicly available dataset for action recognition was initiated at the KTH Royal Institute of Technology in 2004.
279 papers · 2 benchmarks
The Sports-1M dataset consists of over a million videos from YouTube.
164 papers · 2 benchmarks
This paper introduces the pipeline to scale the largest dataset in egocentric vision EPIC-KITCHENS.
162 papers · 6 benchmarks
NTU RGB+D 120 is a large-scale dataset for RGB+D human action recognition, which is collected from 106 distinct subjects and contains more than 114 thousand video samples and 8 million frames.
137 papers · 8 benchmarks
The 20BN-SOMETHING-SOMETHING dataset is a large collection of labeled video clips that show humans performing pre-defined basic actions with everyday objects.
117 papers · 3 benchmarks
AVA (Atomic Visual Actions)
AVA is a project that provides audiovisual annotations of video for improving our understanding of human activity.
113 papers · 7 benchmarks
The Penn Action Dataset contains 2326 video sequences of 15 different actions and human joint annotations for each sequence.
110 papers · 4 benchmarks
The COIN dataset (a large-scale dataset for COmprehensive INstructional video analysis) consists of 11,827 videos related to 180 different tasks in 12 domains (e.g., vehicles, gadgets, etc.) related to our daily life.
105 papers · 2 benchmarks
Comprises 11 hand gesture categories from 29 subjects under 3 illumination conditions.
103 papers · 6 benchmarks
Kinetics-700 is a video dataset of 650,000 clips that covers 700 human action classes.
95 papers · 3 benchmarks
Volleyball is a video action recognition dataset.
80 papers · 3 benchmarks
FineGym is an action recognition dataset build on top of gymnasium videos.
76 papers · 0 benchmarks
HACS (Human Action Clips and Segments)
HACS is a dataset for human action recognition.
75 papers · 2 benchmarks
BABEL is a large dataset with language labels describing the actions being performed in mocap sequences.
72 papers · 1 benchmark
The UTD-MHAD dataset consists of 27 different actions performed by 8 subjects.
62 papers · 2 benchmarks
The MultiTHUMOS dataset contains dense, multilabel, frame-level action annotations for 30 hours across 400 videos in the THUMOS'14 action detection dataset.
58 papers · 3 benchmarks
Assembly101 is a new procedural activity dataset featuring 4321 videos of people assembling and disassembling 101 "take-apart" toy vehicles.
57 papers · 4 benchmarks
Rendered synthetically using a library of standard 3D objects, and tests the ability to recognize compositions of object movements that require long-term reasoning.
51 papers · 3 benchmarks
UAV-Human is a large dataset for human behavior understanding with UAVs.
47 papers · 5 benchmarks
The EPIC-KITCHENS-55 dataset comprises a set of 432 egocentric videos recorded by 32 participants in their kitchens at 60fps with a head mounted camera.
42 papers · 3 benchmarks
EMOTIC (EMOTIons in Context)
The EMOTIC dataset, named after EMOTions In Context, is a database of images with people in real environments, annotated with their apparent emotions.
38 papers · 2 benchmarks
The EgoGesture dataset contains 2,081 RGB-D videos, 24,161 gesture samples and 2,953,224 frames from 50 distinct subjects.
37 papers · 2 benchmarks
The EgoHands dataset contains 48 Google Glass videos of complex, first-person interactions between two people.
34 papers · 0 benchmarks
Contains 68,536 activity instances in 68.8 hours of first and third-person video, making it one of the largest and most diverse egocentric datasets available.
33 papers · 1 benchmark
A new multitask action quality assessment (AQA) dataset, the largest to date, comprising of more than 1600 diving samples; contains detailed annotations for fine-grained action recognition, commentary generation, and estimating the AQA…
33 papers · 2 benchmarks
CholecT50 (Cholecystectomy Action Triplet)
CholecT50 is a dataset of endoscopic videos of laparoscopic cholecystectomy surgery introduced to enable research on fine-grained action recognition in laparoscopic surgery.
32 papers · 5 benchmarks
N-UCLA (Northwestern-UCLA Multiview Action 3D Dataset)
The Multiview 3D event dataset is capture by me and Xiaohan Nie in UCLA.
30 papers · 2 benchmarks
Animal Kingdom is a large and diverse dataset that provides multiple annotated tasks to enable a more thorough understanding of natural animal behaviors.
26 papers · 2 benchmarks
The Drive&Act dataset is a state of the art multi modal benchmark for driver behavior recognition.
26 papers · 1 benchmark
A new video dataset for aerial view concurrent human action detection.
24 papers · 1 benchmark
MMAct is a large-scale dataset for multi/cross modal action understanding.
23 papers · 1 benchmark
BAR (Biased Action Recognition)
Biased Action Recognition (BAR) dataset is a real-world image dataset categorized as six action classes which are biased to distinct places.
22 papers · 1 benchmark
Spatio-temporal action detection is an important and challenging problem in video understanding.
20 papers · 2 benchmarks
The MECCANO dataset is the first dataset of egocentric videos to study human-object interactions in industrial-like settings.
19 papers · 3 benchmarks
The dataset collected at the University of Florence during 2012, has been captured using a Kinect camera.
18 papers · 1 benchmark
17 papers · 3 benchmarks
CholecT45 is a subset of CholecT50 consisting of 45 videos from the Cholec80 dataset.
16 papers · 1 benchmark
HVU (Holistic Video Understanding)
HVU is organized hierarchically in a semantic taxonomy that focuses on multi-label and multi-task video understanding as a comprehensive problem that encompasses the recognition of multiple semantic aspects in the dynamic scene.
16 papers · 0 benchmarks
Jester Gesture Recognition dataset includes 148,092 labeled video clips of humans performing basic, pre-defined hand gestures in front of a laptop camera or webcam.
16 papers · 6 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.