Datasets › MMBench

MMBench

Introduced by YuAn Liu et al. in MMBench: Is Your Multi-modal Model an All-around Player?12 Jul 2023 archive 2025-07-28

MMBench is a multi-modality benchmark. It methodically develops a comprehensive evaluation pipeline, primarily comprised of two elements. The first element is a meticulously curated dataset that surpasses existing similar benchmarks in terms of the number and variety of evaluation questions and abilities. The second element introduces a novel CircularEval strategy and incorporates the use of ChatGPT. This implementation is designed to convert free-form predictions into pre-defined choices, thereby facilitating a more robust evaluation of the model's predictions.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Question Answering MMBench LLaVA-InternLM2-ViT + MoSLoRA GPT-3.5 score 73.8 Mixture-of-Subspaces in Low-Rank Adaptation wutaiqiang/moslora 5 Compare

Papers archive 2025-07-28

4 shown of 4 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 384. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Mixture-of-Subspaces in Low-Rank Adaptation 1 2 16 Jun 2024 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)
CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts 1 1 9 May 2024 ran 10 of 12 samples (2 unverified)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization 1 1 5 Feb 2024 ran 3 of 5 samples (2 unverified; 5 pointer-only for licence)
DreamLLM: Synergistic Multimodal Comprehension and Creation 1 1 20 Sep 2023 ran 3 of 3 samples (0 unverified; 2 pointer-only for licence)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Apache-2.0 license

Modalities archive 2025-07-28

No modality tagged.

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • MMBench

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections