Datasets › UCF101

UCF101 (UCF101 Human Actions dataset)

Introduced by Khurram Soomro et al. in UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild3 Dec 2012 archive 2025-07-28

UCF101 dataset is an extension of UCF50 and consists of 13,320 video clips, which are classified into 101 categories. These 101 categories can be classified into 5 types (Body motion, Human-human interactions, Human-object interactions, Playing musical instruments and Sports). The total length of these video clips is over 27 hours. All the videos are collected from YouTube and have a fixed frame rate of 25 FPS with the resolution of 320 × 240.

Source: Two-stream Collaborative Learning with Spatial-Temporal Attention for Video Classification Image Source: https://www.crcv.ucf.edu/data/UCF101.php

Benchmarks archive 2025-07-28

All 23 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Action Recognition UCF101 FTP-UniFormerV2-L/14 3-fold Accuracy 99.7 Enhancing Video Transformers for Action Understanding... — 91 Compare
Self-Supervised Action Recognition UCF101 VideoMAE V2-g 3-fold Accuracy 99.6 VideoMAE V2: Scaling Video Masked Autoencoders with Dual Masking OpenGVLab/VideoMAEv2 53 Compare
Video Generation UCF-101 W.A.L.T-XL (class-conditional) FVD16 36±2 Photorealistic Video Generation with Diffusion Models — 48 Compare
Zero-Shot Action Recognition UCF101 OTI(ViT-L/14) Top-1 Accuracy 92.8 Orthogonal Temporal Interpolation for Zero-Shot Video Recognition sweetorangezhuyan/mm2023_oti 35 Compare
Video Frame Interpolation UCF101 EMA-VFI PSNR 35.48 Extracting Motion and Appearance via Inter-Frame... mcg-nju/ema-vfi 19 Compare
Prompt Engineering UCF101 PromptKD Harmonic mean 86.10 PromptKD: Unsupervised Prompt Distillation for... zhengli97/promptkd 14 Compare
Self-Supervised Action Recognition UCF101 (finetuned) BraVe:V-FA (TSM-50x2) 3-fold Accuracy 95.7 Broaden Your Views for Self-Supervised Video Learning deepmind/brave 14 Compare
Text-to-Video Generation UCF-101 Snap Video (Zero-shot, 512x288) FVD16 200.2 Snap Video: Scaled Spatiotemporal Transformers for... — 10 Compare
Few Shot Action Recognition UCF101 STRM 1:1 Accuracy 96.8 Spatio-temporal Relation Modeling for Few-shot Action Recognition Anirudh257/strm 7 Compare
Video Generation UCF-101 16 frames, Unconditional, Single GPU TGAN-F Inception Score 22.91 Lower Dimensional Kernels for Video Discriminators HappyBahman/ldvdGAN 7 Compare
Video Generation UCF-101 16 frames, 64x64, Unconditional Video Diffusion Model Inception Score 57 Video Diffusion Models lucidrains/make-a-video-pytorch +4 7 Compare
Video Generation UCF-101 16 frames, 128x128, Unconditional TGANv2 (2020) Inception Score 28.87 Train Sparsely, Generate Densely: Memory-efficient... pfnet-research/tgan2 +1 6 Compare
Action Recognition In Videos UCF101 STM (ImageNet+Kinetics pretrain) 3-fold Accuracy 96.2 STM: SpatioTemporal and Motion Encoding for Action Recognition — 5 Compare
Action Recognition UCF-101 DMC-Net (ResNet-18) 3-fold Accuracy 90.9 DMC-Net: Generating Discriminative Motion Cues for Fast... — 2 Compare
Image Clustering UCF101 TURTLE (CLIP + DINOv2) Accuracy 82.3 Let Go of Your Labels with Unsupervised Transfer mlbio-epfl/turtle 2 Compare
Zero-Shot Learning UCF101 ZLaP Accuracy 76.3 Label Propagation for Zero-shot Classification with... vladan-stojnic/zlap 2 Compare
Action Classification UCF101 Ours Top-1 62.03 SPAct: Self-supervised Privacy Preservation for Action... daveishan/spact 1 Compare
Action Recognition UCF 101 R2+1D-BERT 3-fold Accuracy 98.69 Late Temporal Modeling in 3D CNN Architectures with BERT... artest08/LateTemporalModeling3DCNN +1 1 Compare
Few-Shot Learning UCF101 Variational Prompt Tuning Harmonic mean 79 Bayesian Prompt Learning for Image-Language Model Generalization saic-fi/bayesian-prompt-learning 1 Compare
Human Activity Recognition UCF 101 Label-Ranker Accuracy 89.50% Label Ranker: Self-Aware Preference for Classification... Peihao-Xiang/Label-Ranker 1 Compare
Open Set Action Recognition UCF101-MiTv2 InternVideo AUROC 91.85 InternVideo: General Video Foundation Models via... opengvlab/internvideo +1 1 Compare
Skeleton Based Action Recognition UCF101 Structured Keypoint Pooling Accuracy 87.8 Unified Keypoint-based Action Recognition Framework via... — 1 Compare
Transductive Zero-Shot Classification UCF101 ZLaP Accuracy 77.7 Label Propagation for Zero-shot Classification with... vladan-stojnic/zlap 1 Compare

Papers archive 2025-07-28

30 shown of 228 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,863. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Video-GPT via Next Clip Diffusion 1 1 18 May 2025 not harvested
MMRL++: Parameter-Efficient and Interaction-Aware Representation Learning for Vision-Language Models 1 1 15 May 2025 not harvested
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition 0 1 30 Mar 2025 not harvested
Long-Context Autoregressive Video Modeling with Next-Frame Prediction 1 1 25 Mar 2025 not harvested
LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion Models 1 1 18 Mar 2025 ran 2 of 4 samples (2 unverified)
MMRL: Multi-Modal Representation Learning for Vision-Language Models 1 1 11 Mar 2025 ran 0 of 1 samples (1 unverified)
Label Ranker: Self-Aware Preference for Classification Label Position in Visual Masked Self-Supervised Pre-Trained Model 1 1 3 Mar 2025 not harvested
ACDiT: Interpolating Autoregressive Conditional Modeling and Diffusion Transformer 1 1 10 Dec 2024 ran 1 of 6 samples (5 unverified)
LoCATe-GAT: Modeling Multi-Scale Local Context and Action Relationships for Zero-Shot Action Recognition 1 1 27 Nov 2024 not harvested
LARP: Tokenizing Videos with a Learned Autoregressive Generative Prior 1 1 28 Oct 2024 ran 4 of 13 samples (9 unverified)
HPT++: Hierarchically Prompting Vision-Language Models with Multi-Granularity Knowledge Generation and Improved Structure Modeling 2 1 27 Aug 2024 not harvested
VFIMamba: Video Frame Interpolation with State Space Models 1 1 2 Jul 2024 ran 8 of 8 samples (0 unverified)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation 1 1 13 Jun 2024 ran 5 of 5 samples (0 unverified; 1 pointer-only for licence)
Hierarchical Patch Diffusion Models for High-Resolution Video Generation 0 1 12 Jun 2024 not harvested
Let Go of Your Labels with Unsupervised Transfer 1 1 11 Jun 2024 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
FIFO-Diffusion: Generating Infinite Videos from Text without Training 1 1 19 May 2024 ran 9 of 13 samples (4 unverified; 13 pointer-only for licence)
Leveraging Temporal Contextualization for Video Action Recognition 2 1 15 Apr 2024 ran 2 of 4 samples (2 unverified; 4 pointer-only for licence)
Label Propagation for Zero-shot Classification with Vision-Language Models 1 3 5 Apr 2024 ran 1 of 2 samples (1 unverified)
Prompt Learning via Meta-Regularization 1 1 1 Apr 2024 ran 4 of 7 samples (3 unverified)
Grid Diffusion Models for Text-to-Video Generation 0 1 30 Mar 2024 not harvested
Enhancing Video Transformers for Action Understanding with VLM-aided Training 0 1 24 Mar 2024 not harvested
PromptKD: Unsupervised Prompt Distillation for Vision-Language Models 1 1 5 Mar 2024 not harvested
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis 0 2 22 Feb 2024 not harvested
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
Lumiere: A Space-Time Diffusion Model for Video Generation 1 2 23 Jan 2024 ran 2 of 3 samples (1 unverified)
OmniVec2 - A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning 0 1 1 Jan 2024 not harvested
VideoPoet: A Large Language Model for Zero-Shot Video Generation 0 2 21 Dec 2023 not harvested
EZ-CLIP: Efficient Zeroshot Video Action Recognition 1 1 13 Dec 2023 ran 9 of 13 samples (4 unverified)
Photorealistic Video Generation with Diffusion Models 0 3 11 Dec 2023 not harvested
Learning Hierarchical Prompt with Structured Linguistic Knowledge for Vision-Language Models 2 1 11 Dec 2023 ran 3 of 7 samples (4 unverified)

The full list of 228 is in the JSON twin.

Dataset loaders archive 2025-07-28

4 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • UCF101-skeleton
  • UCF101-MiTv2
  • UCF-101 Zero-shot, 256x256, class-conditional
  • UCF101 (finetuned)
  • UCF 101
  • UCF101
  • UCF-101 16 frames, Unconditional, Single GPU
  • UCF-101 16 frames, 64x64, Unconditional
  • UCF-101 16 frames, 128x128, Unconditional
  • UCF-101

10 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections