Datasets › ActivityNet

ActivityNet

Introduced by Fabian Caba Heilbron et al. in ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding1 Jan 2015 archive 2025-07-28

The ActivityNet dataset contains 200 different types of activities and a total of 849 hours of videos collected from YouTube. ActivityNet is the largest benchmark for temporal activity detection to date in terms of both the number of activity categories and number of videos, making the task particularly challenging. Version 1.3 of the dataset contains 19994 untrimmed videos in total and is divided into three disjoint subsets, training, validation, and testing by a ratio of 2:1:1. On average, each activity category has 137 untrimmed videos. Each video on average has 1.41 activities which are annotated with temporal boundaries. The ground-truth annotations of test videos are not public.

Source: Dynamic Temporal Pyramid Network: A Closer Look at Multi-Scale Modeling for Activity Detection

Benchmarks archive 2025-07-28

All 17 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Temporal Action Localization ActivityNet-1.3 RDFA-S6 (InternVideo2-6B) mAP 42.9 Enhancing Temporal Action Localization: Advanced S6... lsy0882/RDFA-S6 33 Compare
Video Retrieval ActivityNet InternVideo2-6B text-to-video R@1 74.1 InternVideo2: Scaling Foundation Models for Multimodal... opengvlab/internvideo +1 31 Compare
Weakly Supervised Action Localization ActivityNet-1.2 SAL Mean mAP 30.8 Multilevel semantic and adaptive actionness learning for... lizhilin-ustc/SAL 19 Compare
Weakly Supervised Action Localization ActivityNet-1.3 SAL mAP@0.5:0.95 28.8 Multilevel semantic and adaptive actionness learning for... lizhilin-ustc/SAL 17 Compare
Action Recognition ActivityNet Text4Vis (w/ ViT-L) mAP 96.9 Revisiting Classifier: Transferring Vision-Language... whwu95/Cap4Video +4 16 Compare
Zero-Shot Video Retrieval ActivityNet InternVideo2-6B text-to-video R@1 63.2 InternVideo2: Scaling Foundation Models for Multimodal... opengvlab/internvideo +1 12 Compare
Temporal Action Proposal Generation ActivityNet-1.3 AOE-Net AR@100 77.67 AOE-Net: Entities Interactions Modeling with Adaptive... uark-aicv/aoe-net 11 Compare
GZSL Video Classification ActivityNet-GZSL(main) KDA HM 19.67 Boosting Audio-visual Zero-shot Learning with Large... chenhaoxing/KDA 7 Compare
Zero-Shot Action Recognition ActivityNet BIKE Top-1 Accuracy 86.2 Bidirectional Cross-Modal Knowledge Exploration for... whwu95/Cap4Video +4 5 Compare
GZSL Video Classification ActivityNet-GZSL (cls) KDA HM 17.95 Boosting Audio-visual Zero-shot Learning with Large... chenhaoxing/KDA 4 Compare
Action Classification ActivityNet-1.2 W-TALC mAP 93.2 W-TALC: Weakly-supervised Temporal Activity Localization... sujoyp/wtalc-pytorch 3 Compare
Action Classification ActivityNet UniFormerV2-L Top 1 Accuracy 94.7 UniFormerV2: Spatiotemporal Learning by Arming Image... OpenGVLab/UniFormerV2 +1 1 Compare
Action Recognition In Videos ActivityNet LSTM + Pretrained on YT-8M mAP 75.6 YouTube-8M: A Large-Scale Video Classification Benchmark google/youtube-8m +6 1 Compare
Few Shot Temporal Action Localization ActivityNet FS-QAT mIoU 38.5 Few-Shot Temporal Action Localization with Query... sauradip/fewshotQAT 1 Compare
Temporal Action Localization ActivityNet-1.2 DeepMetricLearner mAP IOU@0.5 35.2 Weakly Supervised Temporal Action Localization Using... asrafulashiq/wsad 1 Compare
Visual Question Answering (VQA) ActivityNet BLIP-2 T5 ClipMatch@1 53.39 Open-ended VQA benchmarking of Vision-Language models by... lmb-freiburg/ovqa 1 Compare
Weakly-supervised Temporal Action Localization ActivityNet-1.3 ASM-Loc mAP 25.1 ASM-Loc: Action-aware Segment Modeling for... boheumd/asm-loc 1 Compare

Papers archive 2025-07-28

30 shown of 122 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 807. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Gramian Multimodal Representation Learning and Alignment 2 2 16 Dec 2024 ran 2 of 12 samples (10 unverified)
LoCATe-GAT: Modeling Multi-Scale Local Context and Action Relationships for Zero-Shot Action Recognition 1 1 27 Nov 2024 not harvested
Multilevel semantic and adaptive actionness learning for weakly supervised temporal action localization 1 2 24 Nov 2024 not harvested
Enhancing Temporal Action Localization: Advanced S6 Modeling with Recurrent Mechanism 1 1 18 Jul 2024 not harvested
Weakly supervised temporal action localization with actionness-guided false positive suppression 1 2 15 Apr 2024 not harvested
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection 1 1 7 Apr 2024 not harvested
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 6 22 Mar 2024 not harvested
vid-TLDR: Training Free Token merging for Light-weight Video Transformer 1 2 20 Mar 2024 ran 3 of 4 samples (1 unverified)
Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding 1 1 14 Mar 2024 ran 8 of 13 samples (5 unverified)
Open-ended VQA benchmarking of Vision-Language models by exploiting Classification datasets and their semantic hierarchy 1 1 11 Feb 2024 ran 2 of 2 samples (0 unverified)
Revisiting Foreground and Background Separation in Weakly-supervised Temporal Action Localization: A Clustering-based Approach 1 2 21 Dec 2023 not harvested
RTQ: Rethinking Video-language Understanding Based on Image-text Model 2 1 1 Dec 2023 not harvested
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames 2 1 28 Nov 2023 not harvested
Boosting Audio-visual Zero-shot Learning with Large Language Models 1 2 21 Nov 2023 not harvested
TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding 1 1 29 Oct 2023 ran 11 of 15 samples (4 unverified)
LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment 6 2 3 Oct 2023 ran 7 of 14 samples (7 unverified)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning 1 1 27 Sep 2023 ran 1 of 4 samples (3 unverified)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning 1 1 20 Sep 2023 not harvested
Hyperbolic Audio-visual Zero-shot Learning 0 2 24 Aug 2023 not harvested
UnLoc: A Unified Framework for Video Localization Tasks 1 1 21 Aug 2023 not harvested
Actionness Inconsistency-guided Contrastive Learning for Weakly-supervised Temporal Action Localization 1 2 26 Jun 2023 not harvested
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model 1 1 15 Jun 2023 not harvested
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset 2 1 29 May 2023 ran 15 of 42 samples (27 unverified)
Proposal-Based Multiple Instance Learning for Weakly-Supervised Temporal Action Localization 1 2 29 May 2023 not harvested
VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset 1 1 17 Apr 2023 not harvested
Improve Temporal Action Proposals using Hierarchical Context 0 2 3 Apr 2023 not harvested
Unmasked Teacher: Towards Training-Efficient Video Foundation Models 1 2 28 Mar 2023 ran 3 of 8 samples (5 unverified)
Video-Text as Game Players: Hierarchical Banzhaf Interaction for Cross-Modal Representation Learning 4 1 25 Mar 2023 ran 11 of 16 samples (5 unverified)
DiffusionRet: Generative Text-Video Retrieval with Diffusion Model 4 2 17 Mar 2023 ran 5 of 6 samples (1 unverified)
TriDet: Temporal Action Detection with Relative Boundary Modeling 1 1 13 Mar 2023 ran 5 of 15 samples (10 unverified)

The full list of 122 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • ActivityNet
  • ActivityNet-1.3
  • ActivityNet-1.2
  • ActivityNet-GZSL (cls)
  • ActivityNet-GZSL(main)

5 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections