Datasets › MSVD-QA

MSVD-QA

archive 2025-07-28

The MSVD-QA dataset is a Video Question Answering (VideoQA) dataset. It is based on the existing Microsoft Research Video Description (MSVD) dataset, which consists of about 120K sentences describing more than 2,000 video snippets. In the MSVD-QA dataset, Question-Answer (QA) pairs are generated from these descriptions. The dataset is mainly used in video captioning experiments but due to its large data size, it is also used for VideoQA. It contains 1970 video clips and approximately 50.5K QA pairs.

Benchmarks archive 2025-07-28

All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 59 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 61. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token 1 1 7 Jan 2025 ran 3 of 3 samples (0 unverified; 1 pointer-only for licence)
LinVT: Empower Your Image-level Large Language Model to Understand Videos 1 1 6 Dec 2024 ran 4 of 12 samples (8 unverified; 12 pointer-only for licence)
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 1 17 Nov 2024 ran 5 of 11 samples (6 unverified)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance 1 1 4 Nov 2024 ran 2 of 9 samples (7 unverified; 3 pointer-only for licence)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 1 1 22 Jul 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
Tarsier: Recipes for Training and Evaluating Large Video Description Models 1 1 30 Jun 2024 ran 2 of 2 samples (0 unverified)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding 1 1 13 Jun 2024 ran 6 of 8 samples (2 unverified; 8 pointer-only for licence)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams 1 1 12 Jun 2024 not harvested
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning 1 1 25 Apr 2024 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
Learning to Localize Objects Improves Spatial Reasoning in Visual-LLMs 1 1 11 Apr 2024 not harvested
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding 1 1 8 Apr 2024 ran 7 of 9 samples (2 unverified)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens 2 1 4 Apr 2024 ran 2 of 2 samples (0 unverified)
ST-LLM: Large Language Models Are Effective Temporal Learners 1 1 30 Mar 2024 ran 7 of 11 samples (4 unverified; 3 pointer-only for licence)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM 1 1 27 Mar 2024 ran 3 of 3 samples (0 unverified)
Elysium: Exploring Object-level Perception in Videos via MLLM 1 1 25 Mar 2024 ran 7 of 8 samples (1 unverified; 8 pointer-only for licence)
vid-TLDR: Training Free Token merging for Light-weight Video Transformer 1 1 20 Mar 2024 ran 3 of 4 samples (1 unverified)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
VILA: On Pre-training for Visual Language Models 3 1 12 Dec 2023 not harvested
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 1 28 Nov 2023 ran 7 of 10 samples (3 unverified)
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models 2 2 28 Nov 2023 ran 3 of 4 samples (1 unverified)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection 6 1 16 Nov 2023 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding 4 1 14 Nov 2023 not harvested
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning 1 2 27 Sep 2023 ran 1 of 4 samples (3 unverified)
Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models 1 4 18 Aug 2023 ran 3 of 8 samples (5 unverified; 8 pointer-only for licence)
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding 1 1 31 Jul 2023 not harvested
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models 2 2 9 Jul 2023 not harvested
Lightweight Recurrent Cross-modal Encoder for Video Question Answering 1 1 30 Jun 2023 not harvested
COSA: Concatenated Sample Pretrained Vision-Language Foundation Model 1 1 15 Jun 2023 not harvested
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models 2 1 8 Jun 2023 not harvested
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding 4 1 5 Jun 2023 ran 18 of 25 samples (7 unverified; 9 pointer-only for licence)

The full list of 59 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • MSVD-QA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections