Datasets › MM-Vet

MM-Vet

Introduced by Weihao Yu et al. in MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities4 Aug 2023 archive 2025-07-28

MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Question Answering MM-Vet gemini-2.0-flash-exp GPT-4 score 81.2±0.4 — — 231 Compare
Visual Question Answering (VQA) MM-Vet Lyra-Pro Acc 71.4 Lyra: An Efficient and Speech-Centric Framework for... dvlab-research/Lyra 1 Compare

Papers archive 2025-07-28

30 shown of 147 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 339. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling 1 2 29 Jan 2025 not harvested
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding 1 1 14 Jan 2025 not harvested
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition 1 4 12 Dec 2024 ran 3 of 19 samples (16 unverified)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models 1 2 9 Dec 2024 ran 0 of 15 samples (15 unverified)
ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance 0 1 9 Dec 2024 not harvested
TACO: Learning Multi-modal Action Models with Synthetic Chains-of-Thought-and-Action 1 3 7 Dec 2024 not harvested
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale 1 2 6 Dec 2024 not harvested
LinVT: Empower Your Image-level Large Language Model to Understand Videos 1 1 6 Dec 2024 ran 4 of 12 samples (8 unverified; 12 pointer-only for licence)
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling 1 7 6 Dec 2024 ran 1 of 9 samples (8 unverified)
VisionZip: Longer is Better but Not Necessary in Vision Language Models 1 6 5 Dec 2024 ran 1 of 11 samples (10 unverified)
FlashSloth: Lightning Multimodal Large Language Models via Embedded Visual Compression 1 2 5 Dec 2024 ran 7 of 7 samples (0 unverified; 7 pointer-only for licence)
A Stitch in Time Saves Nine: Small VLM is a Precise Guidance for Accelerating Large VLMs 1 3 4 Dec 2024 ran 4 of 5 samples (1 unverified; 5 pointer-only for licence)
Dynamic-LLaVA: Efficient Multimodal Large Language Models via Dynamic Vision-language Context Sparsification 1 2 1 Dec 2024 ran 13 of 18 samples (5 unverified)
Looking Beyond Text: Reducing Language bias in Large Vision-Language Models via Multimodal Dual-Attention and Soft-Image Guidance 0 2 21 Nov 2024 not harvested
FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression 0 1 21 Nov 2024 not harvested
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation 1 1 12 Nov 2024 not harvested
Aligned Vector Quantization for Edge-Cloud Collabrative Vision-Language Models 0 1 8 Nov 2024 not harvested
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation 1 1 17 Oct 2024 not harvested
Improving Multi-modal Large Language Model through Boosting Vision Capabilities 0 1 17 Oct 2024 not harvested
H2OVL-Mississippi Vision Language Models Technical Report 0 2 17 Oct 2024 not harvested
$γ-$MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models 0 1 17 Oct 2024 not harvested
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models 0 1 16 Oct 2024 not harvested
MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding 1 1 15 Oct 2024 not harvested
MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling 0 2 14 Oct 2024 not harvested
VLFeedback: A Large-Scale AI Feedback Dataset for Large Vision-Language Models Alignment 0 4 12 Oct 2024 ran 5 of 7 samples (2 unverified)
Baichuan-Omni Technical Report 2 1 11 Oct 2024 not harvested
Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training 0 1 10 Oct 2024 not harvested
Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate 1 2 9 Oct 2024 ran 12 of 12 samples (0 unverified)
Gamified crowd-sourcing of high-quality data for visual fine-tuning 0 3 5 Oct 2024 not harvested
MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning 0 6 30 Sep 2024 not harvested

The full list of 147 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-NC 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MM-Vet

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections