Datasets › MVBench

MVBench

Introduced by Kunchang Li et al. in MVBench: A Comprehensive Multi-modal Video Understanding Benchmark28 Nov 2023 archive 2025-07-28

MVBench is a comprehensive Multi-modal Video understanding Benchmark. It was introduced to evaluate the comprehension capabilities of Multi-modal Large Language Models (MLLMs), particularly their temporal understanding in dynamic video tasks. MVBench covers 20 challenging video tasks that cannot be effectively solved with a single frame. It introduces a novel static-to-dynamic method to define these temporal-related tasks. By transforming various static tasks into dynamic ones, it enables the systematic generation of video tasks that require a broad spectrum of temporal skills, ranging from perception to cognition.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

24 shown of 24 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 139. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition 1 1 12 Dec 2024 ran 3 of 19 samples (16 unverified)
LinVT: Empower Your Image-level Large Language Model to Understand Videos 1 1 6 Dec 2024 ran 4 of 12 samples (8 unverified; 12 pointer-only for licence)
TS-LLaVA: Constructing Visual Tokens through Thumbnail-and-Sampling for Training-Free Video Large Language Models 1 1 17 Nov 2024 ran 5 of 11 samples (6 unverified)
PPLLaVA: Varied Video Sequence Understanding With Prompt Guidance 1 1 4 Nov 2024 ran 2 of 9 samples (7 unverified; 3 pointer-only for licence)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning 1 1 25 Oct 2024 ran 3 of 6 samples (3 unverified)
LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding 1 1 22 Oct 2024 ran 6 of 11 samples (5 unverified)
Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution 1 1 19 Sep 2024 ran 3 of 6 samples (3 unverified)
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models 1 1 9 Aug 2024 not harvested
Tarsier: Recipes for Training and Evaluating Large Video Description Models 1 1 30 Jun 2024 ran 2 of 2 samples (0 unverified)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding 1 1 13 Jun 2024 ran 6 of 8 samples (2 unverified; 8 pointer-only for licence)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs 3 1 11 Jun 2024 ran 7 of 17 samples (10 unverified)
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning 1 1 25 Apr 2024 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
ST-LLM: Large Language Models Are Effective Temporal Learners 1 1 30 Mar 2024 ran 7 of 11 samples (4 unverified; 3 pointer-only for licence)
InternVideo2: Scaling Foundation Models for Multimodal Video Understanding 2 2 22 Mar 2024 not harvested
HawkEye: Training Video-Text LLMs for Grounding Text in Videos 1 1 15 Mar 2024 ran 6 of 6 samples (0 unverified; 6 pointer-only for licence)
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models 1 1 8 Feb 2024 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding 2 1 4 Dec 2023 ran 7 of 11 samples (4 unverified)
MVBench: A Comprehensive Multi-modal Video Understanding Benchmark 3 1 28 Nov 2023 ran 7 of 10 samples (3 unverified)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models 2 1 8 Jun 2023 not harvested
Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding 4 1 5 Jun 2023 ran 18 of 25 samples (7 unverified; 9 pointer-only for licence)
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning 4 1 11 May 2023 not harvested
VideoChat: Chat-Centric Video Understanding 1 1 10 May 2023 not harvested
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models 6 1 20 Apr 2023 not harvested
Visual Instruction Tuning 13 1 17 Apr 2023 ran 16 of 51 samples (35 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • MVBench

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections