Datasets › JHMDB

JHMDB (Joint-annotated Human Motion Data Base)

Introduced in Towards Understanding Action Recognition1 Jan 2013 archive 2025-07-28

JHMDB is an action recognition dataset that consists of 960 video sequences belonging to 21 actions. It is a subset of the larger HMDB51 dataset collected from digitized movies and YouTube videos. The dataset contains video and annotation for puppet flow per frame (approximated optimal flow on the person), puppet mask per frame, joint positions per frame, action label per clip and meta label per clip (camera motion, visible body parts, camera viewpoint, number of people, video quality).

Source: Unsupervised Deep Metric Learning via Orthogonality based Probabilistic Loss Image Source: https://arxiv.org/pdf/1712.06316.pdf

Benchmarks archive 2025-07-28

All 9 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 53 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 249. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Scaling Open-Vocabulary Action Detection 1 2 4 Apr 2025 not harvested
Poseidon: A ViT-based Architecture for Multi-Frame Pose Estimation with Adaptive Frame Weighting and Multi-Scale Feature Fusion 1 1 14 Jan 2025 not harvested
Hierarchical Temporal Convolution Network:Towards Privacy-Centric Activity Recognition 1 2 21 Dec 2024 not harvested
Spectrum-guided Multi-granularity Referring Video Object Segmentation 1 1 25 Jul 2023 ran 6 of 9 samples (3 unverified; 9 pointer-only for licence)
SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation 1 2 26 May 2023 not harvested
Kinematic-aware Hierarchical Attention Network for Human Pose Estimation in Videos 1 1 29 Nov 2022 not harvested
Holistic Interaction Transformer Network for Action Detection 1 1 23 Oct 2022 not harvested
Deeply Interleaved Two-Stream Encoder for Referring Video Segmentation 0 1 30 Mar 2022 not harvested
DeciWatch: A Simple Baseline for 10x Efficient 2D and 3D Pose Estimation 1 2 16 Mar 2022 ran 1 of 8 samples (7 unverified)
End-to-End Referring Video Object Segmentation with Multimodal Transformers 2 2 29 Nov 2021 ran 6 of 11 samples (5 unverified)
Hierarchical interaction network for video object segmentation from referring expressions 0 2 22 Nov 2021 not harvested
Do Different Tracking Tasks Require Different Appearance Models? 1 1 5 Jul 2021 ran 6 of 17 samples (11 unverified)
Cross-Modal Progressive Comprehension for Referring Segmentation 1 1 15 May 2021 not harvested
Collaborative Spatial-Temporal Modeling for Language-Queried Video Actor Segmentation 0 1 14 May 2021 not harvested
ClawCraneNet: Leveraging Object-level Relation for Text-based Video Segmentation 0 1 19 Mar 2021 not harvested
Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network 0 1 9 Feb 2021 not harvested
Actor and Action Modular Network for Text-based Video Segmentation 0 1 2 Nov 2020 not harvested
Pose And Joint-Aware Action Recognition 1 1 16 Oct 2020 not harvested
Finding Action Tubes with a Sparse-to-Dense Framework 0 1 30 Aug 2020 not harvested
Polar Relative Positional Encoding for Video-Language Segmentation 0 1 20 Jul 2020 not harvested
Visual-Textual Capsule Routing for Text-Based Video Segmentation 0 1 1 Jun 2020 not harvested
Context Modulated Dynamic Networks for Actor and Action Video Segmentation with Language Queries 0 1 3 Apr 2020 not harvested
Actions as Moving Points 2 1 14 Jan 2020 ran 3 of 9 samples (6 unverified)
You Only Watch Once: A Unified CNN Architecture for Real-Time Spatiotemporal Action Localization 5 2 15 Nov 2019 ran 3 of 12 samples (9 unverified; 2 pointer-only for licence)
Hierarchical Self-Attention Network for Action Localization in Videos 0 2 1 Oct 2019 not harvested
Asymmetric Cross-Guided Attention Network for Actor and Action Video Segmentation From Natural Language Query 1 1 1 Oct 2019 not harvested
Dynamic Kernel Distillation for Efficient Pose Estimation in Videos 0 1 24 Aug 2019 not harvested
Make Skeleton-based Action Recognition Model Smaller, Faster and Better 3 2 23 Jul 2019 not harvested
PA3D: Pose-Action 3D Machine for Video Recognition 0 2 1 Jun 2019 not harvested
TACNet: Transition-Aware Context Network for Spatio-Temporal Action Detection 0 1 31 May 2019 not harvested

The full list of 53 is in the JSON twin.

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • JHMDB (2D poses only)
  • J-HMDB
  • J-HMBD Early Action
  • JHMDB Pose Tracking
  • J-HMDB-21
  • JHMDB

6 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections