Home › Datasets › task › Skeleton Based Action Recognition
Skeleton Based Action Recognition datasets
archive 2025-07-28
30 datasets carry the task tag "Skeleton Based Action Recognition" (the task itself: Skeleton Based Action Recognition), ordered by the archive's paper count. Page 1 of 1: 30 shown of 30. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Skeleton Based Action Recognition datasets 1–30 of 30
UCF101 (UCF101 Human Actions dataset)
UCF101 dataset is an extension of UCF50 and consists of 13,320 video clips, which are classified into 101 categories.
1,863 papers · 23 benchmarks
Kinetics (Kinetics Human Action Video Dataset)
The Kinetics dataset is a large-scale, high-quality dataset for human action recognition in videos.
1,341 papers · 18 benchmarks
The HMDB51 dataset is a large collection of realistic videos from various sources, including movies and web videos.
839 papers · 10 benchmarks
The dataset contains 400 human action classes, with at least 400 video clips for each action.
712 papers · 0 benchmarks
NTU RGB+D is a large-scale dataset for RGB-D human action recognition.
476 papers · 9 benchmarks
JHMDB (Joint-annotated Human Motion Data Base)
JHMDB is an action recognition dataset that consists of 960 video sequences belonging to 21 actions.
249 papers · 9 benchmarks
NTU RGB+D 120 is a large-scale dataset for RGB+D human action recognition, which is collected from 106 distinct subjects and contains more than 114 thousand video samples and 8 million frames.
137 papers · 8 benchmarks
The Penn Action Dataset contains 2326 video sequences of 15 different actions and human joint annotations for each sequence.
110 papers · 4 benchmarks
The PKU-MMD dataset is a large skeleton-based action detection dataset.
72 papers · 4 benchmarks
The UT-Kinect dataset is a dataset for action recognition from depth sequences.
70 papers · 2 benchmarks
The CAD-60 and CAD-120 data sets comprise of RGB-D video sequences of humans performing activities which are recording using the Microsoft Kinect sensor.
65 papers · 1 benchmark
UAV-Human is a large dataset for human behavior understanding with UAVs.
47 papers · 5 benchmarks
SHREC (SHape REtrieval Contest)
The SHREC dataset contains 14 dynamic gestures performed by 28 participants (all participants are right handed) and captured by the Intel RealSense short range depth camera.
32 papers · 8 benchmarks
MSRC-12 (MSRC-12 Kinect Gesture Dataset)
The Microsoft Research Cambridge-12 Kinect gesture data set consists of sequences of human movements, represented as body-part locations, and the associated gesture to be recognized by the system.
31 papers · 2 benchmarks
N-UCLA (Northwestern-UCLA Multiview Action 3D Dataset)
The Multiview 3D event dataset is capture by me and Xiaohan Nie in UCLA.
30 papers · 2 benchmarks
The Gaming 3D Dataset (G3D) focuses on real-time action recognition in a gaming scenario.
28 papers · 2 benchmarks
SBU-Kinect-Interaction dataset version 2.0 comprises of RGB-D video sequences of humans performing interaction activities that are recording using the Microsoft Kinect sensor.
28 papers · 4 benchmarks
The Drive&Act dataset is a state of the art multi modal benchmark for driver behavior recognition.
26 papers · 1 benchmark
The dataset collected at the University of Florence during 2012, has been captured using a Kinect camera.
18 papers · 1 benchmark
First-Person Hand Action Benchmark is a collection of RGB-D video sequences comprised of more than 100K frames of 45 daily hand action categories, involving 26 different objects in several hand configurations.
15 papers · 2 benchmarks
We present a comprehensive framework for egocentric interaction recognition using markerless 3D annotations of two hands manipulating objects.
14 papers · 2 benchmarks
This is a 3D action recognition dataset, also known as 3D Action Pairs dataset.
7 papers · 1 benchmark
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
5 papers · 1 benchmark
TCG (Traffic Control Gesture)
The TCG dataset is used to evaluate Traffic Control Gesture recognition for autonomous driving.
4 papers · 1 benchmark
HDM05 is a MoCap (motion capture) dataset.
3 papers · 1 benchmark
NTU-X is an extended version of popular NTU dataset.
2 papers · 1 benchmark
A curated and 3-D pose-annotated subset of RGB videos sourced from Kinetics-700, a large-scale action dataset.
2 papers · 1 benchmark
A dataset derived from the recently introduced Mimetics dataset.
2 papers · 2 benchmarks
ANUBIS (Skeleton-Based Action Recognition Dataset)
ANUBIS is a large-scale human skeleton dataset containing 80 actions.
1 paper · 0 benchmarks
Metaphorics is a newly introduced non-contextual skeleton action dataset.
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.