Datasets › ActivityNet Captions

ActivityNet Captions

Introduced by Ranjay Krishna et al. in Dense-Captioning Events in Videos1 Jan 2017 archive 2025-07-28

The ActivityNet Captions dataset is built on ActivityNet v1.3 which includes 20k YouTube untrimmed videos with 100k caption annotations. The videos are 120 seconds long on average. Most of the videos contain over 3 annotated events with corresponding start/end time and human-written sentences, which contain 13.5 words on average. The number of videos in train/validation/test split is 10024/4926/5044, respectively.

Source: Bidirectional Attentive Fusion with Context Gating for Dense Video Captioning Image Source: https://cs.stanford.edu/people/ranjaykrishna/densevid/

Benchmarks archive 2025-07-28

All 6 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Dense Video Captioning ActivityNet Captions Vid2Seq METEOR 17 Vid2Seq: Large-Scale Pretraining of a Visual Language... google-research/scenic +2 12 Compare
Natural Language Moment Retrieval ActivityNet Captions GVL (paragraph-level) R@1,IoU=0.5 60.67 Learning Grounded Vision-Language Representation for... zjr2000/gvl 8 Compare
Video Captioning ActivityNet Captions VideoCoCa BLEU4 14.7 VideoCoCa: Video-Text Modeling with Zero-Shot Transfer... — 5 Compare
Live Video Captioning ActivityNet Captions LVC Live Score 20.81 Live Video Captioning gramuah/lvc 1 Compare
Partially Relevant Video Retrieval ActivityNet Captions ms-sl Recall@Sum 140.1 Partially Relevant Video Retrieval HuiGuanLab/ms-sl 1 Compare
Temporal Action Proposal Generation ActivityNet Captions BMT Average F1 60.27 A Better Use of Audio-Visual Cues: Dense Video... v-iashin/video_features +1 1 Compare

Papers archive 2025-07-28

24 shown of 24 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 255. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval 1 1 21 Nov 2024 not harvested
Live Video Captioning 1 1 20 Jun 2024 not harvested
Do You Remember? Dense Video Captioning with Cross-Modal Memory Retrieval 1 1 11 Apr 2024 ran 2 of 3 samples (1 unverified)
UniMD: Towards Unifying Moment Retrieval and Temporal Action Detection 1 1 7 Apr 2024 not harvested
VTimeLLM: Empower LLM to Grasp Video Moments 1 1 30 Nov 2023 ran 5 of 11 samples (6 unverified; 11 pointer-only for licence)
UnLoc: A Unified Framework for Video Localization Tasks 1 2 21 Aug 2023 not harvested
Learning Grounded Vision-Language Representation for Versatile Understanding in Untrimmed Videos 1 3 11 Mar 2023 ran 1 of 12 samples (11 unverified)
Vid2Seq: Large-Scale Pretraining of a Visual Language Model for Dense Video Captioning 3 1 27 Feb 2023 not harvested
VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners 0 1 9 Dec 2022 not harvested
VLTinT: Visual-Linguistic Transformer-in-Transformer for Coherent Video Paragraph Captioning 1 1 28 Nov 2022 not harvested
Partially Relevant Video Retrieval 1 1 26 Aug 2022 not harvested
VLCap: Vision-Language with Contrastive Learning for Coherent Video Paragraph Captioning 1 1 26 Jun 2022 not harvested
End-to-End Dense Video Captioning with Parallel Decoding 2 1 17 Aug 2021 ran 3 of 7 samples (4 unverified)
Global Object Proposals for Improving Multi-Sentence Video Descriptions 1 1 18 Jul 2021 not harvested
TSP: Temporally-Sensitive Pretraining of Video Encoders for Localization Tasks 1 1 23 Nov 2020 not harvested
VLG-Net: Video-Language Graph Matching Network for Video Grounding 1 1 19 Nov 2020 ran 1 of 1 samples (0 unverified)
iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering 0 1 16 Nov 2020 not harvested
COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning 1 1 1 Nov 2020 ran 0 of 6 samples (6 unverified)
Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020 1 1 21 Jun 2020 not harvested
Team RUC_AIM3 Technical Report at Activitynet 2020 Task 2: Exploring Sequential Events Detection for Dense Video Captioning 0 1 14 Jun 2020 not harvested
A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer 2 2 17 May 2020 ran 0 of 7 samples (7 unverified)
MART: Memory-Augmented Recurrent Transformer for Coherent Video Paragraph Captioning 1 1 11 May 2020 ran 1 of 2 samples (1 unverified)
Dense Regression Network for Video Grounding 1 1 7 Apr 2020 not harvested
Multi-modal Dense Video Captioning 4 1 17 Mar 2020 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • ActivityNet Captions

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections