Datasets › MM-Vet v2

MM-Vet v2

Introduced by Weihao Yu et al. in MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities1 Aug 2024 archive 2025-07-28

MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Question Answering MM-Vet v2 gemini-2.0-flash-exp GPT-4 score 77.1±0.1 — — 24 Compare

Papers archive 2025-07-28

14 shown of 14 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 17. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution 8 1 18 Sep 2024 ran 8 of 12 samples (4 unverified)
Claude 3.5 Sonnet Model Card Addendum 0 1 24 Jun 2024 not harvested
How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites 1 1 25 Apr 2024 not harvested
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context 1 1 8 Mar 2024 not harvested
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 1 1 29 Jan 2024 not harvested
Generative Multimodal Models are In-Context Learners 1 1 20 Dec 2023 ran 3 of 4 samples (1 unverified)
Gemini: A Family of Highly Capable Multimodal Models 1 1 19 Dec 2023 not harvested
CogAgent: A Visual Language Model for GUI Agents 3 1 14 Dec 2023 ran 12 of 18 samples (6 unverified)
CogVLM: Visual Expert for Pretrained Language Models 4 1 6 Nov 2023 not harvested
Improved Baselines with Visual Instruction Tuning 9 2 5 Oct 2023 ran 6 of 9 samples (3 unverified; 8 pointer-only for licence)
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond 2 1 24 Aug 2023 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models 2 1 2 Aug 2023 not harvested
MIMIC-IT: Multi-Modal In-Context Instruction Tuning 2 1 8 Jun 2023 not harvested
GPT-4 Technical Report 11 4 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-NC 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MM-Vet v2

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections