{"url":"/dataset/mmvp-vlm","name":"MMVP-VLM","full_name":null,"description_markdown":"The **MMVP-VLM (Multimodal Visual Patterns - Visual Language Models) Benchmark** is specifically designed to systematically evaluate the performance of recent **CLIP-based models** in understanding and processing visual patterns. Let's delve into the details:\r\n\r\n- **Purpose**: The MMVP-VLM Benchmark aims to assess how well CLIP models can match **image-text combinations** that represent distinct visual patterns. It distills a subset of questions from the original MMVP benchmark into simpler language descriptions, categorizing them into different visual patterns. Each visual pattern is represented by **15 text-image pairs**.\r\n\r\n- **Dataset Composition**:\r\n    - **Text-Image Pairs**: The benchmark includes a balanced number of questions for each visual pattern, with each pattern represented by 15 pairs. These pairs are a subset of the MMVP benchmark, supplemented with additional questions for balance.\r\n    - **Visual Patterns**: The questions cover various visual patterns, allowing evaluation of CLIP models' ability to understand and process these patterns.\r\n\r\n- **Insights and Limitations**: By assessing whether CLIP models can accurately match the provided image-text combinations, the MMVP-VLM Benchmark provides insights into the capabilities and limitations of these models.","description_withheld":null,"homepage":"https://tsb0601.github.io/mmvp_blog/","introduced_date":"2024-01-11","introduced_date_note":null,"introduced_by":{"paper":"/paper/eyes-wide-shut-exploring-the-visual","title":"Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs","first_author":"Shengbang Tong","url":null},"license":null,"modalities":[],"tasks":[],"languages":[],"variants":["MMVP-VLM"],"data_loaders":[],"num_papers_in_archive":7,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}