Datasets › ViP-Bench

ViP-Bench (Making Large Multimodal Models Understand Arbitrary Visual Prompts)

Introduced by Mu Cai et al. in ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts1 Dec 2023 archive 2025-07-28

ViP-Bench is a comprehensive benchmark designed to assess the capability of multimodal models in understanding visual prompts across multiple dimensions. It aims to evaluate how well these models interpret various visual prompts, including recognition, OCR, knowledge, math, relationship reasoning, and language generation. ViP-Bench includes a diverse set of 303 images and questions, providing a thorough assessment of visual understanding capabilities at the region level. This benchmark sets a foundation for future research into multimodal models with arbitrary visual prompts.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Question Answering ViP-Bench GPT-4V-turbo-detail:high (Visual Prompt) GPT-4 score (bbox) 60.7 GPT-4 Technical Report openai/evals +10 13 Compare

Papers archive 2025-07-28

9 shown of 9 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 10. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Inst-IT: Boosting Multimodal Instance Understanding via Explicit Visual Prompt Instruction Tuning 1 2 4 Dec 2024 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Making Large Language Models Better Data Creators 1 1 31 Oct 2023 not harvested
Improved Baselines with Visual Instruction Tuning 9 2 5 Oct 2023 ran 6 of 9 samples (3 unverified; 8 pointer-only for licence)
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond 2 2 24 Aug 2023 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest 3 1 7 Jul 2023 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic 1 1 27 Jun 2023 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Kosmos-2: Grounding Multimodal Large Language Models to the World 2 1 26 Jun 2023 not harvested
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning 4 1 11 May 2023 not harvested
GPT-4 Technical Report 11 2 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

apache-2.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • ViP-Bench

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections