{"url":"/dataset/vlm2-bench","name":"VLM2-Bench","full_name":"VLM²-Bench","description_markdown":"## **VLM²-Bench: Benchmarking Vision-Language Models on Visual Cue Matching**\r\n\r\n### **Description**\r\nVLM²-Bench is the **first** comprehensive benchmark designed to evaluate vision-language models' (VLMs) ability to **visually link matching cues** across multi-image sequences and videos. The benchmark consists of **9 subtasks** with over **3,000 test cases**, focusing on fundamental **visual linking** capabilities that humans use daily. A key example is **identifying the same person across different photos without prior knowledge of their identity**.\r\n\r\nThrough extensive evaluation of **eight open-source VLMs and GPT-4o** using various prompting techniques, we uncover **significant challenges** in **visual cue linking**. Even the best-performing model, **GPT-4o**, falls **34.80% below human-level performance**. Our analysis highlights critical areas for improvement:\r\n1. **Enhancing core visual understanding** with reduced reliance on prior knowledge.\r\n2. **Better integration of language reasoning** within visual tasks.\r\n3. **Developing training approaches** that improve independent **visual relationship inference**.\r\n\r\n### **Dataset Characteristics**\r\n- **Size:** 3,000+ test cases\r\n- **Modalities:** Text, image, video\r\n- **Question Types:** True/False, multiple-choice, numerical, open-ended\r\n- **Generation Process:** Semi-automated with human verification\r\n- **Structure:** Organized into three primary categories:\r\n  - **General Cue (GC):** Evaluates visual element tracking and matching.\r\n  - **Object-centric Cue (OC):** Focuses on object comparison, counting, and grouping.\r\n  - **Person-centric Cue (PC):** Measures the ability to compare, count, group, and describe individuals across frames.\r\n\r\n### **Potential Use Cases**\r\n- **Benchmarking vision-language models (VLMs)** for real-world multi-modal reasoning.\r\n- **Evaluating visual linking abilities and spatial awareness** in large models.\r\n- **Analyzing weaknesses in object permanence and relational inference**.\r\n- **Providing insights for improving next-generation vision-language architectures**.\r\n\r\n### **Paper & Code**\r\n📄 **Paper:** [VLM²-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues](https://arxiv.org/abs/2502.12084)  \r\n📂 **Code Repository:** [GitHub - vlm2-bench/VLM2-Bench](https://github.com/vlm2-bench/VLM2-Bench)\r\n\r\n### **BibTeX Citation**\r\n```\r\n@misc{zhang2025vlm2benchcloserlookvlms,\r\n      title={VLM$^2$-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues}, \r\n      author={Jianshu Zhang and Dongyu Yao and Renjie Pi and Paul Pu Liang and Yi R. Fung},\r\n      year={2025},\r\n      eprint={2502.12084},\r\n      archivePrefix={arXiv},\r\n      primaryClass={cs.CL},\r\n      url={https://arxiv.org/abs/2502.12084}\r\n}\r\n```","description_withheld":null,"homepage":"https://vlm2-bench.github.io/","introduced_date":"2025-02-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/vlm-2-bench-a-closer-look-at-how-well-vlms","title":"VLM2-Bench: A Closer Look at How Well VLMs Implicitly Link Explicit Matching Visual Cues","first_author":"Jianshu Zhang","url":null},"license":{"name":"CC BY-NC 4.0 License","url":"https://creativecommons.org/licenses/by-nc/4.0/"},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"},{"name":"Video Question Answering","url":"/task/video-question-answering","datasets_with_task":"/datasets/task/video-question-answering"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["VLM2-Bench"],"data_loaders":[],"num_papers_in_archive":9,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/visual-question-answering-vqa-on-vlm2-bench","task":"Visual Question Answering (VQA)","dataset_variant":"VLM2-Bench","rows":9,"metrics":["GC-mat","GC-trk","OC-cpr","OC-cnt","OC-grp","PC-cpr","PC-cnt","PC-grp","PC-VID","Average Score on VLM2-bench (9 subtasks)"],"first_row_in_archive_order":{"model":"GPT-4o","paper":"/paper/gpt-4o-system-card","metrics":{"Average Score on VLM2-bench (9 subtasks)":"60.36","GC-mat":"37.45","GC-trk":"39.27","OC-cnt":"80.62","OC-cpr":"74.17","OC-grp":"57.50","PC-VID":"66.75","PC-cnt":"90.50","PC-cpr":"50.00","PC-grp":"47.00"},"code_links":[]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/qwen2-5-vl-technical-report","title":"Qwen2.5-VL Technical Report","date":"2025-02-19","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":2,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/expanding-performance-boundaries-of-open","title":"Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling","date":"2024-12-06","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":1,"samples_unverified":8,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/gpt-4o-system-card","title":"GPT-4o System Card","date":"2024-10-25","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/video-instruction-tuning-with-synthetic-data","title":"Video Instruction Tuning With Synthetic Data","date":"2024-10-03","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/qwen2-vl-enhancing-vision-language-model-s","title":"Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution","date":"2024-09-18","rows_on_this_dataset":1,"code_links":8,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":8,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/mplug-owl3-towards-long-image-sequence","title":"mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models","date":"2024-08-09","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/llava-onevision-easy-visual-task-transfer","title":"LLaVA-OneVision: Easy Visual Task Transfer","date":"2024-08-06","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/long-context-transfer-from-language-to-vision","title":"Long Context Transfer from Language to Vision","date":"2024-06-24","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":5,"samples_unverified":0,"pointer_only_for_licence":5,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":4,"samples_harvested":29,"samples_ran":16,"samples_unverified":13,"pointer_only_for_licence":5,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}