{"url":"/dataset/vnbench","name":"VNBench","full_name":null,"description_markdown":"VNBench is a comprehensive benchmark suite for video generative models, which evaluates video generation quality across specific, hierarchical, and disentangled dimensions, each with tailored prompts and evaluation methods. The suite includes 16 dimensions for evaluating Text-to-Video (T2V) models, such as subject consistency, motion smoothness, and overall consistency. VNBench also supports evaluating Image-to-Video (I2V) models and has recently introduced VBench-Long for evaluating longer videos¹. It's designed to align with human perceptions and provide valuable insights for future developments in video generation.\r\n\r\n# Diversity of “Needle” Types\r\nEdit: Using artificially added subtitles as the \"needle\". These subtitles are embedded in video frames to simulate the scenario of finding specific textual information in a video.\r\n\r\nInsert: Using images as the \"needle\". These images are inserted as static segments between video frames to assess the model's ability to recognize and remember static images in a video.\r\n\r\nLevel Classification: Classified into two levels based on image recognizability. The first level uses common objects (e.g., fruit images), while the second level uses more challenging landmark or object images, increasing the task's difficulty.\r\n\r\n# Diversity of Video “Haystack”\r\nTemporal Distribution: The video \"haystack\" used by VNBench comes from different data sources, with video durations ranging from 10 seconds to 180 seconds. This covers short, medium, and long video lengths to evaluate the model's adaptability to different video lengths.\r\n\r\nContent Coverage: The video content includes various scenes, ensuring the evaluation's broadness and the diversity of video sources.\r\n\r\n# Diversity of Queries\r\nRetrieval Task: Requires the model to retrieve specific \"needles\" from videos, assessing the model's fine-grained understanding and information extraction ability.\r\n\r\nOrdering Task: Requires the model to identify and order the timestamps of all inserted \"needles\" in the video, assessing the model's understanding of video temporal dynamics and event sequences.\r\n\r\nCounting Task: Requires the model to count the occurrences of specific objects in the video, including recognizing and tracking repetitive patterns within and across frames. This assesses the model's understanding of spatial and temporal dimensions.\r\n\r\n(1) Vchitect/VBench: [CVPR2024 Highlight] VBench - GitHub. https://github.com/Vchitect/VBench.\r\n(2) VBench: Comprehensive Benchmark Suite for Video Generative Models. https://arxiv.org/abs/2311.17982.\r\n(3) Linux中vdbench的安装与使用_vdbench参数详解-CSDN博客. https://blog.csdn.net/SweeNeil/article/details/95338293.\r\n(4) Cinebench 2024 Downloads - Maxon. https://www.maxon.net/zh/downloads/cinebench-2024-downloads.\r\n(5) undefined. https://doi.org/10.48550/arXiv.2311.17982.","description_withheld":null,"homepage":"https://github.com/joez17/VideoNIAH","introduced_date":"2024-06-13","introduced_date_note":null,"introduced_by":{"paper":"/paper/needle-in-a-video-haystack-a-scalable","title":"Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs","first_author":"Zijia Zhao","url":null},"license":null,"modalities":[],"tasks":[{"name":"Zero-Shot Video Question Answer","url":"/task/zeroshot-video-question-answer","datasets_with_task":"/datasets/task/zeroshot-video-question-answer"}],"languages":[],"variants":["VNBench"],"data_loaders":[],"num_papers_in_archive":11,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/zero-shot-video-question-answer-on-vnbench","task":"Zero-Shot Video Question Answer","dataset_variant":"VNBench","rows":9,"metrics":["Accuracy"],"first_row_in_archive_order":{"model":"BIMBA-LLaVA-Qwen2-7B","paper":"/paper/bimba-selective-scan-compression-for-long","metrics":{"Accuracy":"77.88"},"code_links":[{"title":"md-mohaiminul/BIMBA","url":"https://github.com/md-mohaiminul/BIMBA"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/bimba-selective-scan-compression-for-long","title":"BIMBA: Selective-Scan Compression for Long-Range Video Question Answering","date":"2025-03-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/qwen2-vl-enhancing-vision-language-model-s","title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","date":"2024-09-18","rows_on_this_dataset":1,"code_links":8,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":8,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/llava-onevision-easy-visual-task-transfer","title":"LLaVA-OneVision: Easy Visual Task Transfer","date":"2024-08-06","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/llava-next-interleave-tackling-multi-image","title":"LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models","date":"2024-07-10","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":2,"samples_unverified":2,"pointer_only_for_licence":4,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/videollama-2-advancing-spatial-temporal","title":"VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs","date":"2024-06-11","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":17,"samples_ran":7,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/gemini-1-5-unlocking-multimodal-understanding","title":"Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context","date":"2024-03-08","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/video-chatgpt-towards-detailed-video","title":"Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models","date":"2023-06-08","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/videochat-chat-centric-video-understanding","title":"VideoChat: Chat-Centric Video Understanding","date":"2023-05-10","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":3,"samples_harvested":33,"samples_ran":17,"samples_unverified":16,"pointer_only_for_licence":4,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}