Datasets › Video-MME
Video-MME
Video-MME stands for Video Multi-Modal Evaluation. It is the first-ever comprehensive evaluation benchmark specifically designed for Multi-modal Large Language Models (MLLMs) in video analysis¹. This benchmark is significant because it addresses the need for a high-quality assessment of MLLMs' performance in processing sequential visual data, which has been less explored compared to their capabilities in static image understanding.
The Video-MME benchmark is characterized by its: 1. Diversity in video types, covering 6 primary visual domains with 30 subfields for broad scenario generalizability. 2. Duration in the temporal dimension, including short-, medium-, and long-term videos ranging from 11 seconds to 1 hour, to assess robust contextual dynamics. 3. Breadth in data modalities, integrating multi-modal inputs such as video frames, subtitles, and audios. 4. Quality in annotations, with rigorous manual labeling by expert annotators for precise and reliable model assessment¹.
The benchmark includes 900 videos totaling 256 hours, manually selected and annotated, resulting in 2,700 question-answer pairs. It has been used to evaluate various state-of-the-art MLLMs, including the GPT-4 series and Gemini 1.5 Pro, as well as open-source image and video models¹. The findings from Video-MME highlight the need for further improvements in handling longer sequences and multi-modal data, which is crucial for the advancement of MLLMs¹.
(1) [2405.21075] Video-MME: The First-Ever Comprehensive Evaluation .... https://arxiv.org/abs/2405.21075. (2) Video-MME. https://video-mme.github.io/home_page.html. (3) Video-MME: Welcome. https://video-mme.github.io/. (4) undefined. https://doi.org/10.48550/arXiv.2405.21075.
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Zero-Shot Video Question Answer | Video-MME | Gemini 1.5 Pro Accuracy (%) 81.3 | Gemini 1.5: Unlocking multimodal understanding across... | dlvuldet/primevul | 11 | Compare |
| Zero-Shot Video Question Answer | Video-MME (w/o subs) | Video-RAG (based on LLaVA-Video) Accuracy (%) 77.4 | Video-RAG: Visually-aligned Retrieval-Augmented Long... | leon1207/video-rag-master | 9 | Compare |
Papers archive 2025-07-28
9 shown of 9 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 152. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- Video-MME
- Video-MME (w/o subs)
2 variant names, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections