Datasets › HMDB51

HMDB51

Introduced in HMDB: A large video database for human motion recognition1 Jan 2011 archive 2025-07-28

The HMDB51 dataset is a large collection of realistic videos from various sources, including movies and web videos. The dataset is composed of 6,766 video clips from 51 action categories (such as “jump”, “kiss” and “laugh”), with each category containing at least 101 clips. The original evaluation scheme uses three different training/testing splits. In each split, each action class has 70 clips for training and 30 clips for testing. The average accuracy over these three splits is used to measure the final performance.

Source: Action Recognition with Trajectory-Pooled Deep-Convolutional Descriptors Image Source: https://serre-lab.clps.brown.edu/resource/hmdb-a-large-human-motion-database

Benchmarks archive 2025-07-28

All 10 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Action Recognition HMDB-51 VideoMAE V2-g Average accuracy of 3 splits 88.7 VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking OpenGVLab/VideoMAEv2 77 Compare
Self-Supervised Action Recognition HMDB51 MVD (ViT-B) Top-1 Accuracy 79.7 Masked Video Distillation: Rethinking Masked Feature... ruiwang2021/mvd +3 48 Compare
Zero-Shot Action Recognition HMDB51 MOV (ViT-L/14) Top-1 Accuracy 64.7 Multimodal Open-Vocabulary Video Classification via... — 29 Compare
Self-Supervised Action Recognition HMDB51 (finetuned) BraVe:V-FA (TSM-50x2) Top-1 Accuracy 77.8 Broaden Your Views for Self-Supervised Video Learning deepmind/brave 14 Compare
Few Shot Action Recognition HMDB51 STRM 1:1 Accuracy 77.3 Spatio-temporal Relation Modeling for Few-shot Action Recognition Anirudh257/strm 7 Compare
Skeleton Based Action Recognition HMDB51 Structured Keypoint Pooling Accuracy 70.9 Unified Keypoint-based Action Recognition Framework via... — 2 Compare
Action Classification HMDB51 DualPath w/ ViT-B/16 MLPs. Acc@1 75.6 Dual-path Adaptation from Image to Video Transformers park-jungin/dualpath 1 Compare
Action Recognition In Videos HMDB-51 STM (ImageNet+Kinetics pretrain) Average accuracy of 3 splits 72.2 STM: SpatioTemporal and Motion Encoding for Action Recognition — 1 Compare
Action Recognition HMDB51 MSQNet Accuracy 93.25 Actor-agnostic Multi-label Action Recognition with... mondalanindya/msqnet 1 Compare
Human Activity Recognition HMDB51 Label-Ranker Accuracy 61.18% Label Ranker: Self-Aware Preference for Classification... Peihao-Xiang/Label-Ranker 1 Compare

Papers archive 2025-07-28

30 shown of 125 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 839. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Label Ranker: Self-Aware Preference for Classification Label Position in Visual Masked Self-Supervised Pre-Trained Model 1 1 3 Mar 2025 not harvested
IoT-Based Real-Time Medical-Related Human Activity Recognition Using Skeletons and Multi-Stage Deep Learning for Healthcare 1 1 13 Jan 2025 not harvested
DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification 1 1 1 Jan 2025 not harvested
LoCATe-GAT: Modeling Multi-Scale Local Context and Action Relationships for Zero-Shot Action Recognition 1 1 27 Nov 2024 not harvested
Leveraging Temporal Contextualization for Video Action Recognition 2 1 15 Apr 2024 ran 2 of 4 samples (2 unverified; 4 pointer-only for licence)
OST: Refining Text Knowledge with Optimal Spatio-Temporal Descriptor for General Video Recognition 1 1 30 Nov 2023 ran 1 of 2 samples (1 unverified)
Asymmetric Masked Distillation for Pre-Training Small Foundation Models 0 1 6 Nov 2023 not harvested
ZeroI2V: Zero-Cost Adaptation of Pre-trained Transformers from Image to Video 2 1 2 Oct 2023 ran 1 of 3 samples (2 unverified)
Orthogonal Temporal Interpolation for Zero-Shot Video Recognition 1 1 14 Aug 2023 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Actor-agnostic Multi-label Action Recognition with Multi-modal Query 1 2 20 Jul 2023 not harvested
Alternating Gradient Descent and Mixture-of-Experts for Integrated Multimodal Perception 0 1 10 May 2023 not harvested
Synthetic Sample Selection for Generalized Zero-Shot Learning 0 1 6 Apr 2023 not harvested
VicTR: Video-conditioned Text Representations for Activity Recognition 0 1 5 Apr 2023 not harvested
VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking 1 1 29 Mar 2023 ran 2 of 6 samples (4 unverified)
Unified Keypoint-based Action Recognition Framework via Structured Keypoint Pooling 0 1 27 Mar 2023 not harvested
Dual-path Adaptation from Image to Video Transformers 1 1 17 Mar 2023 not harvested
MAtch, eXpand and Improve: Unsupervised Finetuning for Zero-Shot Action Recognition with Language Knowledge 1 1 15 Mar 2023 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
Bidirectional Cross-Modal Knowledge Exploration for Video Recognition with Pre-trained Vision-Language Models 5 2 31 Dec 2022 not harvested
Similarity Contrastive Estimation for Image and Video Soft Contrastive Self-Supervised Learning 2 1 21 Dec 2022 not harvested
VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners 0 1 9 Dec 2022 not harvested
Masked Video Distillation: Rethinking Masked Feature Modeling for Self-supervised Video Representation Learning 4 1 8 Dec 2022 not harvested
XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning 1 2 25 Nov 2022 not harvested
EVEREST: Efficient Masked Video Autoencoder by Removing Redundant Spatiotemporal Tokens 2 1 19 Nov 2022 ran 3 of 6 samples (3 unverified; 6 pointer-only for licence)
Masked Motion Encoding for Self-Supervised Video Representation Learning 2 1 12 Oct 2022 not harvested
Expanding Language-Image Pretrained Models for General Video Recognition 2 1 4 Aug 2022 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
Multimodal Open-Vocabulary Video Classification via Pre-Trained Vision and Language Models 0 2 15 Jul 2022 not harvested
Revisiting Classifier: Transferring Vision-Language Models for Video Recognition 5 1 4 Jul 2022 not harvested
SLIC: Self-Supervised Learning with Iterative Clustering for Human Action Videos 1 1 25 Jun 2022 not harvested
Learn2Augment: Learning to Composite Videos for Data Augmentation in Action Recognition 0 1 9 Jun 2022 not harvested
Cross-modal Representation Learning for Zero-shot Action Recognition 0 1 3 May 2022 not harvested

The full list of 125 is in the JSON twin.

Dataset loaders archive 2025-07-28

5 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • HMDB51-skeleton
  • HMDB51 (finetuned)
  • HMDB51
  • HMDB-51

4 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections