Datasets › BenchLMM

BenchLMM (BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models)

Introduced by Rizhao Cai et al. in BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models5 Dec 2023 archive 2025-07-28

Large Multimodal Models (LMMs) such as GPT-4V and LLaVA have shown remarkable capabilities in visual reasoning with common image styles. However, their robustness against diverse style shifts, crucial for practical applications, remains largely unexplored. In this paper, we propose a new benchmark, BenchLMM, to assess the robustness of LMMs against three different styles: artistic image style, imaging sensor style, and application style, where each style has five sub-styles. Utilizing BenchLMM, we comprehensively evaluate state-of-the-art LMMs and reveal: 1) LMMs generally suffer performance degradation when working with other styles; 2) An LMM performs better than another model in common style does not guarantee its superior performance in other styles; 3) LMMs' reasoning capability can be enhanced by prompting LMMs to predict the style first, based on which we propose a versatile and training-free method for improving LMMs; 4) An intelligent LMM is expected to interpret the causes of its errors when facing stylistic variations. We hope that our benchmark and analysis can shed new light on developing more intelligent and versatile LMMs.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Question Answering BenchLMM GPT-4V GPT-3.5 score 58.37 GPT-4 Technical Report openai/evals +10 10 Compare

Papers archive 2025-07-28

8 shown of 8 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 12. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models 1 1 13 Nov 2023 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning 2 1 14 Oct 2023 ran 3 of 3 samples (0 unverified; 1 pointer-only for licence)
Improved Baselines with Visual Instruction Tuning 9 1 5 Oct 2023 ran 6 of 9 samples (3 unverified; 8 pointer-only for licence)
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning 4 2 11 May 2023 not harvested
Otter: A Multi-Modal Model with In-Context Instruction Tuning 1 1 5 May 2023 not harvested
MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models 6 1 20 Apr 2023 not harvested
Visual Instruction Tuning 13 2 17 Apr 2023 ran 16 of 51 samples (35 unverified)
GPT-4 Technical Report 11 1 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Apache 2.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • BenchLMM

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections