Datasets › MSRVTT-QA

MSRVTT-QA

archive 2025-07-28

The MSR-VTT-QA dataset is a benchmark for the task of Visual Question Answering (VQA) on the MSR-VTT (Microsoft Research Video to Text) dataset. The MSR-VTT-QA benchmark is used to evaluate models on their ability to answer questions based on these videos. It's part of the tasks that this dataset is used for, along with Video Retrieval, Video Captioning, Zero-Shot Video Question Answering, Zero-Shot Video Retrieval, and Text-to-Video Generation.

Benchmarks archive 2025-07-28

All 5 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Question Answering (VQA) MSRVTT-QA VLAB Accuracy 0.496 VLAB: Enhancing Video Language Pre-training by Feature... — 34 Compare
Zero-Shot Video Question Answer MSRVTT-QA Flash-VStream Accuracy 72.4 Flash-VStream: Memory-Based Real-Time Understanding for... IVGSZ/Flash-VStream 30 Compare
Video Question Answering MSRVTT-QA Mirasol3B Accuracy 50.42 Mirasol3B: A Multimodal Autoregressive model for... — 14 Compare
Visual Question Answering MSRVTT-QA Aurora (ours, r=64) Aurora (ours, r=64) Test Accuracy 44.8 — — 4 Compare
Zero-Shot Learning MSRVTT-QA HiTeA Accuracy 21.7 HiTeA: Hierarchical Temporal-Aware Video-Language Pre-training — 1 Compare

Papers archive 2025-07-28

30 shown of 66 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 66. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
LLaVA-Mini: Efficient Image and Video Large Multimodal Models with One Vision Token 1 1 7 Jan 2025 ran 3 of 3 samples (0 unverified; 1 pointer-only for licence)
LinVT: Empower Your Image-level Large Language Model to Understand Videos 1 1 6 Dec 2024 ran 4 of 12 samples (8 unverified; 12 pointer-only for licence)
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 1 17 Nov 2024 ran 5 of 11 samples (6 unverified)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance 1 1 4 Nov 2024 ran 2 of 9 samples (7 unverified; 3 pointer-only for licence)
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models 1 1 22 Jul 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
Tarsier: Recipes for Training and Evaluating Large Video Description Models 1 1 30 Jun 2024 ran 2 of 2 samples (0 unverified)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding 1 1 13 Jun 2024 ran 6 of 8 samples (2 unverified; 8 pointer-only for licence)
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams 1 1 12 Jun 2024 not harvested
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning 1 1 25 Apr 2024 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
MA-LMM: Memory-Augmented Large Multimodal Model for Long-Term Video Understanding 1 1 8 Apr 2024 ran 7 of 9 samples (2 unverified)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens 2 1 4 Apr 2024 ran 2 of 2 samples (0 unverified)
ST-LLM: Large Language Models Are Effective Temporal Learners 1 1 30 Mar 2024 ran 7 of 11 samples (4 unverified; 3 pointer-only for licence)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM 1 1 27 Mar 2024 ran 3 of 3 samples (0 unverified)
Elysium: Exploring Object-level Perception in Videos via MLLM 1 1 25 Mar 2024 ran 7 of 8 samples (1 unverified; 8 pointer-only for licence)
vid-TLDR: Training Free Token merging for Light-weight Video Transformer 1 1 20 Mar 2024 ran 3 of 4 samples (1 unverified)
CAT: Enhancing Multimodal Large Language Model to Answer Questions in Dynamic Audio-Visual Scenarios 1 1 7 Mar 2024 ran 6 of 12 samples (6 unverified)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
Shot2Story20K: A New Benchmark for Comprehensive Understanding of Multi-shot Videos 1 1 16 Dec 2023 ran 6 of 7 samples (1 unverified; 7 pointer-only for licence)
Vista-LLaMA: Reliable Video Narrator via Equal Distance to Visual Tokens 0 1 12 Dec 2023 not harvested
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 1 28 Nov 2023 ran 7 of 10 samples (3 unverified)
LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models 2 2 28 Nov 2023 ran 3 of 4 samples (1 unverified)
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection 6 1 16 Nov 2023 ran 4 of 7 samples (3 unverified; 1 pointer-only for licence)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding 4 1 14 Nov 2023 not harvested
Mirasol3B: A Multimodal Autoregressive model for time-aligned and contextual modalities 0 1 9 Nov 2023 not harvested
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning 1 2 27 Sep 2023 ran 1 of 4 samples (3 unverified)
Open-vocabulary Video Question Answering: A New Benchmark for Evaluating the Generalizability of Video Question Answering Models 1 3 18 Aug 2023 ran 3 of 8 samples (5 unverified; 8 pointer-only for licence)
OmniDataComposer: A Unified Data Structure for Multimodal Data Fusion and Infinite Data Generation 1 1 8 Aug 2023 not harvested
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding 1 1 31 Jul 2023 not harvested
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models 2 3 9 Jul 2023 not harvested
Lightweight Recurrent Cross-modal Encoder for Video Question Answering 1 1 30 Jun 2023 not harvested

The full list of 66 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • MSRVTT-QA

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections