Datasets › MM-Vet › Papers, page 2

MM-Vet

Papers archive 2025-07-28

papers with a benchmark row: 147 · with a code link: 108 · where Syntology ran a sample: 69 (57 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument) Syntology

Show: all papers with a benchmark rowonly where code ran (69 of 147 with a benchmark row: 57 with a run with no instrument failure, 12 where every run was a failure of Syntology's instrument)

Page 2 of 2: papers 101 to 147 of 147 with a leaderboard row on this dataset's benchmarks, newest first by the archive's date (ties by slug; undated papers last).

The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 339. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record.

PaperCodeResultsDateSamples run Syntology
TinyLLaVA: A Framework of Small-scale Large Multimodal Models 2 1 22 Feb 2024 official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 4 where Syntology's instrument failed) · 4 unverified (1 pointer-only for licence)
CoLLaVO: Crayon Large Language and Vision mOdel 1 1 17 Feb 2024 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified
SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models 1 1 8 Feb 2024 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (5 pointer-only for licence)
From Training-Free to Adaptive: Empirical Insights into MLLMs' Understanding of Detection Information 0 1 31 Jan 2024 not harvested
MouSi: Poly-Visual-Expert Vision-Language Models 1 1 30 Jan 2024 not harvested
InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model 1 1 29 Jan 2024 not harvested
MoE-LLaVA: Mixture of Experts for Large Vision-Language Models 3 1 29 Jan 2024 official: harvested, nothing ran · 0 ran · 2 unverified (2 pointer-only for licence)
Small Language Model Meets with Reinforced Vision Vocabulary 0 1 23 Jan 2024 not harvested
COCO is "ALL'' You Need for Visual Instruction Fine-tuning 0 1 17 Jan 2024 not harvested
CaMML: Context-Aware Multimodal Learner for Large Models 1 1 6 Jan 2024 not harvested
LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model 1 1 4 Jan 2024 official (archive's flag): 6 ran · 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 1 violated, 2 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (9 pointer-only for licence)
V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs 1 1 21 Dec 2023 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 2 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified
Generative Multimodal Models are In-Context Learners 1 1 20 Dec 2023 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
Gemini: A Family of Highly Capable Multimodal Models 1 1 19 Dec 2023 not harvested
Silkie: Preference Distillation for Large Visual Language Models 0 1 17 Dec 2023 not harvested
CogAgent: A Visual Language Model for GUI Agents 3 1 14 Dec 2023 official (archive's flag): 12 ran · 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 6 unverified (1 pointer-only for licence)
Hallucination Augmented Contrastive Learning for Multimodal Large Language Model 1 1 12 Dec 2023 official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 3 unverified (5 pointer-only for licence)
VILA: On Pre-training for Visual Language Models 3 1 12 Dec 2023 not harvested
Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models 1 1 11 Dec 2023 not harvested
OneLLM: One Framework to Align All Modalities with Language 1 1 6 Dec 2023 official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 1 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (4 pointer-only for licence)
Merlin:Empowering Multimodal LLMs with Foresight Minds 0 1 30 Nov 2023 not harvested
ShareGPT4V: Improving Large Multi-Modal Models with Better Captions 1 2 21 Nov 2023 not harvested
Video-LLaVA: Learning United Visual Representation by Alignment Before Projection 6 1 16 Nov 2023 community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (2 pointer-only for licence)
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models 1 1 13 Nov 2023 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence)
To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning 2 1 13 Nov 2023 not harvested
Volcano: Mitigating Multimodal Hallucination through Self-Feedback Guided Revision 1 2 13 Nov 2023 official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 1 violated, 3 with no contract checked; 4 where Syntology's instrument failed) · 1 unverified (10 pointer-only for licence)
InfMLLM: A Unified Framework for Visual-Language Tasks 2 1 12 Nov 2023 not harvested
LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents 1 2 9 Nov 2023 official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration 2 1 7 Nov 2023 official: no sample here; runs from other or unrecorded repositories · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 2 unverified (3 pointer-only for licence)
OtterHD: A High-Resolution Multi-modality Model 1 1 7 Nov 2023 official (archive's flag): 13 ran · 13 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (3 pointer-only for licence)
CogVLM: Visual Expert for Pretrained Language Models 4 2 6 Nov 2023 not harvested
Improved Baselines with Visual Instruction Tuning 9 2 5 Oct 2023 6 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 3 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 3 unverified (8 pointer-only for licence)
DreamLLM: Synergistic Multimodal Comprehension and Creation 1 1 20 Sep 2023 official (archive's flag): 1 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence)
An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models 1 1 18 Sep 2023 official (archive's flag): 1 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence)
TextBind: Multi-turn Interleaved Multimodal Instruction-following in the Wild 1 1 14 Sep 2023 not harvested
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond 2 2 24 Aug 2023 official: harvested for another paper · 0 ran · 2 unverified (2 pointer-only for licence)
StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data 1 1 20 Aug 2023 official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 1 violated, 3 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)
OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models 2 2 2 Aug 2023 not harvested
Emu: Generative Pretraining in Multimodality 2 1 11 Jul 2023 official: harvested for another paper · 0 ran · 2 unverified
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning 4 1 26 Jun 2023 not harvested
MIMIC-IT: Multi-Modal In-Context Instruction Tuning 2 2 8 Jun 2023 not harvested
LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model 3 1 28 Apr 2023 official: not harvested · 0 ran · 1 unverified (1 pointer-only for licence)
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models 6 2 20 Apr 2023 not harvested
MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action 1 2 20 Mar 2023 official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
GPT-4 Technical Report 11 5 15 Mar 2023 community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models 17 1 30 Jan 2023 community repositories only · 4 ran (of which 4 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified; every one of the 4 samples that ran constructed an object rather than computing a result (1 pointer-only for licence)