Home › Datasets › task › Temporal Action Localization
Temporal Action Localization datasets
archive 2025-07-28
42 datasets carry the task tag "Temporal Action Localization" (the task itself: Temporal Action Localization), ordered by the archive's paper count. Page 1 of 1: 42 shown of 42. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Temporal Action Localization datasets 1–42 of 42
UCF101 (UCF101 Human Actions dataset)
UCF101 dataset is an extension of UCF50 and consists of 13,320 video clips, which are classified into 101 categories.
1,863 papers · 23 benchmarks
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
The HMDB51 dataset is a large collection of realistic videos from various sources, including movies and web videos.
839 papers · 10 benchmarks
The ActivityNet dataset contains 200 different types of activities and a total of 849 hours of videos collected from YouTube.
807 papers · 17 benchmarks
The MPII Human Pose Dataset for single person pose estimation is composed of about 25K images of which 15K are training samples, 3K are validation samples and 7K are testing samples (which labels are withheld by the authors).
495 papers · 4 benchmarks
The Charades dataset is composed of 9,848 videos of daily indoors activities with an average length of 30 seconds, involving interactions with 46 objects classes in 15 types of indoor scenes and containing a vocabulary of 30 verbs leading…
428 papers · 6 benchmarks
The THUMOS14 (THUMOS 2014) dataset is a large-scale video dataset that includes 1,010 videos for validation and 1,574 videos for testing from 20 classes.
318 papers · 18 benchmarks
The efforts to create a non-trivial and publicly available dataset for action recognition was initiated at the KTH Royal Institute of Technology in 2004.
279 papers · 2 benchmarks
This paper introduces the pipeline to scale the largest dataset in egocentric vision EPIC-KITCHENS.
162 papers · 6 benchmarks
The COIN dataset (a large-scale dataset for COmprehensive INstructional video analysis) consists of 11,827 videos related to 180 different tasks in 12 domains (e.g., vehicles, gadgets, etc.) related to our daily life.
105 papers · 2 benchmarks
Kinetics-700 is a video dataset of 650,000 clips that covers 700 human action classes.
95 papers · 3 benchmarks
FineGym is an action recognition dataset build on top of gymnasium videos.
76 papers · 0 benchmarks
HACS (Human Action Clips and Segments)
HACS is a dataset for human action recognition.
75 papers · 2 benchmarks
BABEL is a large dataset with language labels describing the actions being performed in mocap sequences.
72 papers · 1 benchmark
The UTD-MHAD dataset consists of 27 different actions performed by 8 subjects.
62 papers · 2 benchmarks
The MultiTHUMOS dataset contains dense, multilabel, frame-level action annotations for 30 hours across 400 videos in the THUMOS'14 action detection dataset.
58 papers · 3 benchmarks
CrossTask dataset contains instructional videos, collected for 83 different tasks.
54 papers · 1 benchmark
Ego4D is a massive-scale egocentric video dataset and benchmark suite.
32 papers · 6 benchmarks
A three million frame, multi-view, furniture assembly video dataset that includes depth, atomic actions, object segmentation, and human pose.
25 papers · 1 benchmark
FineAction contains 103K temporal instances of 106 action categories, annotated in 17K untrimmed videos.
24 papers · 3 benchmarks
A new large-scale dataset for understanding human motions, poses, and actions in a variety of realistic events, especially crowd & complex events.
19 papers · 1 benchmark
The dataset collected at the University of Florence during 2012, has been captured using a Kinect camera.
18 papers · 1 benchmark
HVU (Holistic Video Understanding)
HVU is organized hierarchically in a semantic taxonomy that focuses on multi-label and multi-task video understanding as a comprehensive problem that encompasses the recognition of multiple semantic aspects in the dynamic scene.
16 papers · 0 benchmarks
MUSES (MUlti-Shot EventS)
MUSES is a large-scale dataset for temporal event (action) localization.
12 papers · 1 benchmark
Perception Test is a benchmark designed to evaluate the perception and reasoning skills of multimodal models.
10 papers · 3 benchmarks
Includes egocentric videos containing hands in the wild.
7 papers · 0 benchmarks
This is a 3D action recognition dataset, also known as 3D Action Pairs dataset.
7 papers · 1 benchmark
The TUM Kitchen dataset is an action recognition dataset that contains 20 video sequences captured by 4 cameras with overlapping views.
7 papers · 0 benchmarks
A dataset which provides detailed annotations for activity recognition.
5 papers · 1 benchmark
WEAR (WEAR: An Outdoor Sports Dataset for Wearable and Egocentric Activity Recognition)
WEAR is an outdoor sports dataset for both vision- and inertial-based human activity recognition (HAR).
5 papers · 0 benchmarks
TinyVIRAT contains natural low-resolution activities.
4 papers · 0 benchmarks
A novel dataset that represents complex conversational interactions between two individuals via 3D pose.
3 papers · 0 benchmarks
The Composable activities dataset consists of 693 videos that contain activities in 16 classes performed by 14 actors.
3 papers · 0 benchmarks
DECADE is a large-scale dataset of ego-centric videos from a dog's perspective as well as her corresponding movements.
3 papers · 0 benchmarks
A dataset for benchmarking action recognition algorithms in natural environments, while making use of 3D information.
3 papers · 0 benchmarks
OREBA (Objectively Recognizing Eating Behavior and Associated Intake)
The OREBA dataset aims to provide a comprehensive multi-sensor recording of communal intake occasions for researchers interested in automatic detection of intake gestures.
3 papers · 0 benchmarks
UAV-GESTURE is a dataset for UAV control and gesture recognition.
3 papers · 0 benchmarks
MCAD (Multi-Camera Action Dataset)
Designed to evaluate the open view classification problem under the surveillance environment.
2 papers · 0 benchmarks
RISE is a large-scale video dataset for Recognizing Industrial Smoke Emissions.
2 papers · 0 benchmarks
A curated and 3-D pose-annotated subset of RGB videos sourced from Kinetics-700, a large-scale action dataset.
2 papers · 1 benchmark
Metaphorics is a newly introduced non-contextual skeleton action dataset.
1 paper · 0 benchmarks
WhenAct (Temporal Human Action Localization in Lifestyle Vlogs)
We consider the task of temporal human action localization in lifestyle vlogs.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.