Datasets › ActivityNet-QA

ActivityNet-QA

Introduced by Zhou Yu et al. in ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering archive 2025-07-28

The ActivityNet-QA dataset contains 58,000 human-annotated QA pairs on 5,800 videos derived from the popular ActivityNet dataset. The dataset provides a benchmark for testing the performance of VideoQA models on long-term spatio-temporal reasoning.

Source: ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Video Question Answering ActivityNet-QA GPT-2 + CLIP-14 + CLIP-multilingual (Zero-Shot) Accuracy 61.2 Composing Ensembles of Pre-trained Models via Iterative Consensus — 36 Compare
Zero-Shot Video Question Answer ActivityNet-QA Tarsier (34B) Accuracy 61.6 Tarsier: Recipes for Training and Evaluating Large Video... bytedance/tarsier 28 Compare

Papers archive 2025-07-28

30 shown of 42 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 146. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token 1 1 7 Jan 2025 ran 3 of 3 samples (0 unverified; 1 pointer-only for licence)
LinVT: Empower Your Image-level Large Language Model to Understand Videos 1 1 6 Dec 2024 ran 4 of 12 samples (8 unverified; 12 pointer-only for licence)
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 1 17 Nov 2024 ran 5 of 11 samples (6 unverified)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance 1 1 4 Nov 2024 ran 2 of 9 samples (7 unverified; 3 pointer-only for licence)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 1 1 22 Jul 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
Tarsier: Recipes for Training and Evaluating Large Video Description Models 1 1 30 Jun 2024 ran 2 of 2 samples (0 unverified)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding 1 1 13 Jun 2024 ran 6 of 8 samples (2 unverified; 8 pointer-only for licence)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams 1 1 12 Jun 2024 not harvested
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning 1 1 25 Apr 2024 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs 1 2 11 Apr 2024 not harvested
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding 1 1 8 Apr 2024 ran 7 of 9 samples (2 unverified)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens 2 1 4 Apr 2024 ran 2 of 2 samples (0 unverified)
ST-LLM: Large Language Models Are Effective Temporal Learners 1 1 30 Mar 2024 ran 7 of 11 samples (4 unverified; 3 pointer-only for licence)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM 1 1 27 Mar 2024 ran 3 of 3 samples (0 unverified)
Elysium: Exploring Object-level Perception in Videos via MLLM 1 1 25 Mar 2024 ran 7 of 8 samples (1 unverified; 8 pointer-only for licence)
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios 1 1 7 Mar 2024 ran 6 of 12 samples (6 unverified)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 2 28 Nov 2023 ran 7 of 10 samples (3 unverified)
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models 2 4 28 Nov 2023 ran 3 of 4 samples (1 unverified)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection 6 2 16 Nov 2023 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding 4 3 14 Nov 2023 not harvested
Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities 0 1 9 Nov 2023 not harvested
TESTA: Temporal-Spatial Token Aggregation for Long-form Video-Language Understanding 1 1 29 Oct 2023 ran 11 of 15 samples (4 unverified)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning 1 2 27 Sep 2023 ran 1 of 4 samples (3 unverified)
Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models 1 3 18 Aug 2023 ran 3 of 8 samples (5 unverified; 8 pointer-only for licence)
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding 1 2 31 Jul 2023 not harvested
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model 1 1 15 Jun 2023 not harvested
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models 2 2 8 Jun 2023 not harvested
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding 4 1 5 Jun 2023 ran 18 of 25 samples (7 unverified; 9 pointer-only for licence)
VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset 2 1 29 May 2023 ran 15 of 42 samples (27 unverified)

The full list of 42 is in the JSON twin.

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • ActivityNet-QA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections