Datasets › VNBench
VNBench
VNBench is a comprehensive benchmark suite for video generative models, which evaluates video generation quality across specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. The suite includes 16 dimensions for evaluating Text-to-Video (T2V) models, such as subject consistency, motion smoothness, and overall consistency. VNBench also supports evaluating Image-to-Video (I2V) models and has recently introduced VBench-Long for evaluating longer videos¹. It's designed to align with human perceptions and provide valuable insights for future developments in video generation.
Diversity of “Needle” Types
Edit: Using artificially added subtitles as the "needle". These subtitles are embedded in video frames to simulate the scenario of finding specific textual information in a video.
Insert: Using images as the "needle". These images are inserted as static segments between video frames to assess the model's ability to recognize and remember static images in a video.
Level Classification: Classified into two levels based on image recognizability. The first level uses common objects (e.g., fruit images), while the second level uses more challenging landmark or object images, increasing the task's difficulty.
Diversity of Video “Haystack”
Temporal Distribution: The video "haystack" used by VNBench comes from different data sources, with video durations ranging from 10 seconds to 180 seconds. This covers short, medium, and long video lengths to evaluate the model's adaptability to different video lengths.
Content Coverage: The video content includes various scenes, ensuring the evaluation's broadness and the diversity of video sources.
Diversity of Queries
Retrieval Task: Requires the model to retrieve specific "needles" from videos, assessing the model's fine-grained understanding and information extraction ability.
Ordering Task: Requires the model to identify and order the timestamps of all inserted "needles" in the video, assessing the model's understanding of video temporal dynamics and event sequences.
Counting Task: Requires the model to count the occurrences of specific objects in the video, including recognizing and tracking repetitive patterns within and across frames. This assesses the model's understanding of spatial and temporal dimensions.
(1) Vchitect/VBench: [CVPR2024 Highlight] VBench - GitHub. https://github.com/Vchitect/VBench. (2) VBench: Comprehensive Benchmark Suite for Video Generative Models. https://arxiv.org/abs/2311.17982. (3) Linux中vdbench的安装与使用_vdbench参数详解-CSDN博客. https://blog.csdn.net/SweeNeil/article/details/95338293. (4) Cinebench 2024 Downloads - Maxon. https://www.maxon.net/zh/downloads/cinebench-2024-downloads. (5) undefined. https://doi.org/10.48550/arXiv.2311.17982.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Zero-Shot Video Question Answer | VNBench | BIMBA-LLaVA-Qwen2-7B Accuracy 77.88 | BIMBA: Selective-Scan Compression for Long-Range Video... | md-mohaiminul/BIMBA | 9 | Compare |
Papers archive 2025-07-28
8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 11. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| BIMBA: Selective-Scan Compression for Long-Range Video Question Answering | 1 | 1 | 12 Mar 2025 | not harvested |
| Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution | 8 | 1 | 18 Sep 2024 | ran 8 of 12 samples (4 unverified) |
| LLaVA-OneVision: Easy Visual Task Transfer | 2 | 2 | 6 Aug 2024 | not harvested |
| LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models | 3 | 1 | 10 Jul 2024 | ran 2 of 4 samples (2 unverified; 4 pointer-only for licence) |
| VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs | 3 | 1 | 11 Jun 2024 | ran 7 of 17 samples (10 unverified) |
| Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context | 1 | 1 | 8 Mar 2024 | not harvested |
| Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models | 2 | 1 | 8 Jun 2023 | not harvested |
| VideoChat: Chat-Centric Video Understanding | 1 | 1 | 10 May 2023 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- VNBench
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections