{"url":"/task/vcgbench-diverse","name":"VCGBench-Diverse","slug":"vcgbench-diverse","description_markdown":"Recognizing the limited diversity in existing video conversation benchmarks, we introduce VCGBench-Diverse to comprehensively evaluate the generalization ability of video LMMs. While VCG-Bench provides an extensive evaluation protocol, it is limited to videos from the ActivityNet200 dataset. Our benchmark comprises a total of 877 videos, 18 broad video categories and 4,354 QA pairs, ensuring a robust evaluation framework.\r\n\r\nThe evaluation is computed over five different aspects: \r\n\r\n1. Correctness of information \r\n\r\n2. Detail orientation \r\n\r\n3. Contextual understanding \r\n\r\n4. Temporal understanding \r\n\r\n5. Consistency. \r\n\r\nAdditionally, VCGBench-Diverse provides a breakdown of performance across three key aspects: \r\n\r\n1. Dense video captioning, which assesses the ability to generate detailed and accurate descriptions of the video content, \r\n\r\n2. Spatial understanding, which evaluates the capability to understand and describe the spatial relationships and settings within the video \r\n\r\n3. Reasoning, which tests the adeptness in inferring and explaining causal relationships and actions within the video.","categories":[{"name":"Computer Vision","url":"/area/computer-vision"}],"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","slug_source":"archive_url"},"counts":{"papers_tagged":5,"papers_with_code":5,"benchmarks":1,"benchmark_tables_in_archive":1,"benchmark_tables_shown":1,"benchmark_tables_withheld_as_spam":0,"benchmark_definition":"a leaderboard table with at least one row; benchmark_tables_shown also counts the zero-row tables; benchmark_tables_in_archive adds the tables withheld as spam","datasets":1,"subtasks":0,"parent_tasks":0},"benchmarks":[{"leaderboard":"/sota/vcgbench-diverse-on-videoinstruct","slug":"vcgbench-diverse-on-videoinstruct","dataset":"VideoInstruct","dataset_url":"/dataset/videoinstruct","rows_in_archive":6,"metrics":["mean","Correctness of Information","Detail Orientation","Contextual Understanding","Temporal Understanding","Consistency","Dense Captioning","Spatial Understanding","Reasoning"],"first_row_in_archive_order":{"model":"VideoGPT+","paper_title":"VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding","paper_url":"/paper/videogpt-integrating-image-and-video-encoders","paper_date":"2024-06-13","arxiv_id":"2406.09418","code_links":[{"title":"mbzuai-oryx/videogpt-plus","url":"https://github.com/mbzuai-oryx/videogpt-plus"}],"syntology":{"n":8,"n_ran":6,"n_unverified":2,"n_pointer_only":8}}}],"datasets":[{"url":"/dataset/videoinstruct","name":"VideoInstruct","full_name":"Video Instruction Dataset","num_papers_in_archive":30}],"subtasks":[],"parent_tasks":[],"papers":{"order":"repositories listed in the archive (desc), then date (desc); the archive holds no stars","population":"papers tagged with this task that list at least one repository in the archive","shown":5,"of":5,"tagged_in_all":5,"items":[{"url":"/paper/chat-univi-unified-visual-representation","title":"Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding","date":"2023-11-14","arxiv_id":"2311.08046","repositories_listed":4,"syntology":null},{"url":"/paper/mvbench-a-comprehensive-multi-modal-video","title":"MVBench: A Comprehensive Multi-modal Video Understanding Benchmark","date":"2023-11-28","arxiv_id":"2311.17005","repositories_listed":3,"syntology":{"n":10,"n_ran":7,"n_unverified":3,"n_pointer_only":0}},{"url":"/paper/video-chatgpt-towards-detailed-video","title":"Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models","date":"2023-06-08","arxiv_id":"2306.05424","repositories_listed":2,"syntology":null},{"url":"/paper/videogpt-integrating-image-and-video-encoders","title":"VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding","date":"2024-06-13","arxiv_id":"2406.09418","repositories_listed":1,"syntology":{"n":8,"n_ran":6,"n_unverified":2,"n_pointer_only":8}},{"url":"/paper/vtimellm-empower-llm-to-grasp-video-moments","title":"VTimeLLM: Empower LLM to Grasp Video Moments","date":"2023-11-30","arxiv_id":"2311.18445","repositories_listed":1,"syntology":{"n":11,"n_ran":5,"n_unverified":6,"n_pointer_only":11}}],"syntology_records":3,"syntology_note":"a paper without a record is not a recorded non-run: it may lack an arXiv id or simply be absent from the graph layer"},"description_links":{"kept":0,"unwrapped_to_text":0,"bare_urls_linked":0,"relative_images_dropped":0,"rule":"internal links are kept only when the target slug exists in the catalog"},"syntology":{"read_at":"2026-09-24T18:15:14+00:00","claim":"Per-sample execution status on synthesized fixtures ('ran N of M samples'); not a correctness claim and not a ranking signal.","status_vocabulary":{"ran_honours":"ran, honoured the contract we drafted","ran_violates":"ran, violated the contract we drafted","ran_draft_wrong":"ran; our contract draft was wrong, not the code","ran_fixture":"ran; our fixture could not drive it","ran":"ran on a synthesized input","unverified":"unverified (harvested, no recorded run)"}},"not_shown":{"libraries":"the archive has no per-task library table","trend_sparklines":"the Trend column of the benchmarks table was a rendered image; it is not in the archive","social_and_latest_sorts":"stars and social signals are not in the archive"}}