{"url":"/dataset/cv-bench","name":"CV-Bench","full_name":"Cambrian Vision-Centric Benchmark","description_markdown":"The Cambrian Vision-Centric Benchmark (CV-Bench) is designed to address the limitations of existing vision-centric benchmarks by providing a comprehensive evaluation framework for multimodal large language models (MLLMs). With 2,638 manually-inspected examples, CV-Bench significantly surpasses other vision-centric MLLM benchmarks, offering 3.5 times more examples than RealWorldQA and 8.8 times more than MMVP.\r\n\r\n**Motivation and Content Summary:**\r\n\r\nCV-Bench repurposes standard vision benchmarks such as ADE20K, COCO, and Omni3D to assess models on classic vision tasks within a multimodal context. Leveraging the rich ground truth annotations from these benchmarks, natural language questions are formulated to probe the fundamental 2D and 3D understanding of models.\r\n\r\n**Potential Use Cases:**\r\n\r\n- Evaluating the spatial relationship and object counting capabilities of models (2D understanding).\r\n- Assessing the depth order and relative distance understanding of models (3D understanding).\r\n- Benchmarking the performance of multimodal models in both vision-specific and cross-modal tasks.\r\n\r\n**Dataset Characteristics:**\r\n\r\n- **2D Understanding Tasks:**\r\n  - **Spatial Relationship:** Determine the relative position of an object with respect to the anchor object, considering left-right or top-bottom relationships.\r\n  - **Object Count:** Determine the number of instances present in the image.\r\n\r\n- **3D Understanding Tasks:**\r\n  - **Depth Order:** Determine which of the two distinct objects is closer to the camera.\r\n  - **Relative Distance:** Determine which of the two distinct objects is closer to the anchor object.\r\n\r\n| Type | Task                 | Description                                                                 | Sources        | # Samples |\r\n|------|----------------------|-----------------------------------------------------------------------------|----------------|-----------|\r\n| 2D   | Spatial Relationship | Determine the relative position of an object w.r.t. the anchor object.       | ADE20K, COCO   | 650       |\r\n| 2D   | Object Count         | Determine the number of instances present in the image.                     | ADE20K, COCO   | 788       |\r\n| 3D   | Depth Order          | Determine which of the two distinct objects is closer to the camera.         | Omni3D         | 600       |\r\n| 3D   | Relative Distance    | Determine which of the two distinct objects is closer to the anchor object.  | Omni3D         | 600       |\r\n\r\n**Curation Process:**\r\n\r\nQuestions for each task are programmatically constructed and then manually inspected to ensure clarity and accuracy. Any unclear, ambiguous, or erroneous questions are removed to maintain the benchmark's reliability.","description_withheld":null,"homepage":"https://cambrian-mllm.github.io","introduced_date":"2024-06-24","introduced_date_note":null,"introduced_by":{"paper":"/paper/cambrian-1-a-fully-open-vision-centric","title":"Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs","first_author":"Shengbang Tong","url":null},"license":{"name":"Apache-2.0 license","url":null},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Visual Question Answering (VQA)","url":"/task/visual-question-answering","datasets_with_task":"/datasets/task/visual-question-answering"}],"languages":[],"variants":["CV-Bench"],"data_loaders":[],"num_papers_in_archive":37,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-25T09:33:49+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}