Datasets › NTU RGB+D

NTU RGB+D

Introduced by Amir Shahroudy et al. in NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis1 Jan 2016 archive 2025-07-28

NTU RGB+D is a large-scale dataset for RGB-D human action recognition. It involves 56,880 samples of 60 action classes collected from 40 subjects. The actions can be generally divided into three categories: 40 daily actions (e.g., drinking, eating, reading), nine health-related actions (e.g., sneezing, staggering, falling down), and 11 mutual actions (e.g., punching, kicking, hugging). These actions take place under 17 different scene conditions corresponding to 17 video sequences (i.e., S001–S017). The actions were captured using three cameras with different horizontal imaging viewpoints, namely, −45∘,0∘, and +45∘. Multi-modality information is provided for action characterization, including depth maps, 3D skeleton joint position, RGB frames, and infrared sequences. The performance evaluation is performed by a cross-subject test that split the 40 subjects into training and test groups, and by a cross-view test that employed one camera (+45∘) for testing, and the other two cameras for training.

Source: Action Recognition for Depth Video using Multi-view Dynamic Images

Benchmarks archive 2025-07-28

All 9 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Skeleton Based Action Recognition NTU RGB+D Hulk(Finetune, ViT-L) Accuracy (CS) 94.3 Hulk: A Universal Knowledge Translator for Human-Centric Tasks opengvlab/humanbench +1 135 Compare
Action Recognition NTU RGB+D DSCNet (RGB + Pose) Accuracy (CS) 97.4 A Dense-Sparse Complementary Network for Human Action... Maxchengqin/DSCNet 28 Compare
Zero Shot Skeletal Action Recognition NTU RGB+D TDSM Accuracy (12 unseen classes) 56.03 TDSM: Triplet Diffusion for Skeleton-Text Matching in... KAIST-VICLab/TDSM 9 Compare
3D Action Recognition NTU RGB+D Kinet Cross Subject Accuracy 92.3 No Pain, Big Gain: Classify Dynamic Point Cloud... jx-zhong-for-academic-purpose/kinet 5 Compare
Human Interaction Recognition NTU RGB+D SkateFormer Accuracy (Cross-Subject) 97.1 SkateFormer: Skeletal-Temporal Transformer for Human... KAIST-VICLab/SkateFormer 5 Compare
Generalized Zero Shot skeletal action recognition NTU RGB+D MSF-GZSSAR Harmonic Mean (5 unseen classes) 68.83 Multi-Semantic Fusion Model for Generalized Zero-Shot... EHZ9NIWI7/MSF-GZSSAR 4 Compare
Human action generation NTU RGB+D Kinetic-GAN FID (CS) 3.618 Generative Adversarial Graph Convolutional Networks for... degardinbruno/kinetic-gan 3 Compare
Pose Prediction Filtered NTU RGB+D PISEP^2 (L1 norm) MSE 0.1210 PISEP^2: Pseudo Image Sequence Evolution based 3D Pose Prediction — 3 Compare
Action Recognition In Videos NTU RGB+D 2D-3D-Softargmax (RGB only) Accuracy (CS) 85.5 2D/3D Pose Estimation and Action Recognition using... dluvizon/deephar +1 1 Compare

Papers archive 2025-07-28

30 shown of 162 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 476. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Zero-shot Skeleton-based Action Recognition with Prototype-guided Feature Alignment 1 1 1 Jul 2025 not harvested
DSTSA-GCN: Advancing Skeleton-Based Gesture Recognition with Semantic-Aware Spatio-Temporal Topology Modeling 1 2 21 Jan 2025 not harvested
MSA-GCN: Exploiting Multi-Scale Temporal Dynamics With Adaptive Graph Convolution for Skeleton-Based Action Recognition 0 1 19 Dec 2024 not harvested
USDRL: Unified Skeleton-Based Dense Representation Learning with Multi-Grained Feature Decorrelation 1 1 12 Dec 2024 not harvested
Revealing Key Details to See Differences: A Novel Prototypical Perspective for Skeleton-based Action Recognition 1 1 28 Nov 2024 not harvested
TDSM: Triplet Diffusion for Skeleton-Text Matching in Zero-Shot Action Recognition 1 1 16 Nov 2024 not harvested
Joint Mixing Data Augmentation for Skeleton-based Action Recognition 1 1 13 Oct 2024 not harvested
CHASE: Learning Convex Hull Adaptive Shift for Skeleton-based Multi-Entity Action Recognition 1 1 9 Oct 2024 ran 2 of 2 samples (0 unverified)
Action Recognition for Privacy-Preserving Ambient Assisted Living 1 1 15 Aug 2024 not harvested
EPAM-Net: An Efficient Pose-driven Attention-guided Multimodal Network for Video Action Recognition 1 1 10 Aug 2024 not harvested
Joint-Partition Group Attention for skeleton-based action recognition 1 2 30 Jul 2024 not harvested
Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition 1 1 22 Jul 2024 ran 9 of 10 samples (1 unverified; 10 pointer-only for licence)
SA-DVAE: Improving Zero-Shot Skeleton-Based Action Recognition by Disentangled Variational Autoencoders 1 4 18 Jul 2024 not harvested
Shap-Mix: Shapley Value Guided Mixing for Long-Tailed Skeleton Based Action Recognition 1 1 17 Jul 2024 not harvested
Frequency Guidance Matters: Skeletal Action Recognition by Frequency-Aware Mixed Transformer 1 1 17 Jul 2024 not harvested
Part-aware Unified Representation of Language and Skeleton for Zero-shot Action Recognition 1 1 19 Jun 2024 ran 3 of 3 samples (0 unverified)
Fine-Grained Side Information Guided Dual-Prompts for Zero-Shot Skeleton Action Recognition 0 1 11 Apr 2024 not harvested
DeGCN: Deformable Graph Convolutional Networks for Skeleton-Based Action Recognition 1 1 25 Mar 2024 not harvested
SkateFormer: Skeletal-Temporal Transformer for Human Action Recognition 1 2 14 Mar 2024 ran 3 of 3 samples (0 unverified)
AutoGCN -- Towards Generic Human Activity Recognition with Neural Architecture Search 1 1 2 Feb 2024 not harvested
Explore Human Parsing Modality for Action Recognition 1 1 4 Jan 2024 ran 7 of 10 samples (3 unverified; 10 pointer-only for licence)
MaskCLR: Attention-Guided Contrastive Learning for Robust Action Representation Learning 0 1 1 Jan 2024 not harvested
BlockGCN: Redefine Topology Awareness for Skeleton-Based Action Recognition 1 1 1 Jan 2024 not harvested
A Dense-Sparse Complementary Network for Human Action Recognition based on RGB and Skeleton Modalities 1 1 28 Dec 2023 not harvested
DVANet: Disentangling View and Action Features for Multi-View Action Recognition 1 1 10 Dec 2023 not harvested
STEP CATFormer: Spatial-Temporal Effective Body-Part Cross Attention Transformer for Skeleton-based Action Recognition 1 1 6 Dec 2023 ran 10 of 13 samples (3 unverified)
Hulk: A Universal Knowledge Translator for Human-Centric Tasks 2 2 4 Dec 2023 ran 12 of 24 samples (12 unverified)
Just Add π! Pose Induced Video Transformers for Understanding Activities of Daily Living 1 2 30 Nov 2023 not harvested
Multi-Semantic Fusion Model for Generalized Zero-Shot Skeleton-Based Action Recognition 1 1 18 Sep 2023 not harvested
B2C-AFM: Bi-Directional Co-Temporal and Cross-Spatial Attention Fusion Model for Human Action Recognition 1 1 30 Aug 2023 not harvested

The full list of 162 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom (research-only, non-commercial, attribution)

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • NTU RGB+D
  • Filtered NTU RGB+D

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections