Datasets › MMBench
MMBench
MMBench is a multi-modality benchmark. It methodically develops a comprehensive evaluation pipeline, primarily comprised of two elements. The first element is a meticulously curated dataset that surpasses existing similar benchmarks in terms of the number and variety of evaluation questions and abilities. The second element introduces a novel CircularEval strategy and incorporates the use of ChatGPT. This implementation is designed to convert free-form predictions into pre-defined choices, thereby facilitating a more robust evaluation of the model's predictions.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Visual Question Answering | MMBench | LLaVA-InternLM2-ViT + MoSLoRA GPT-3.5 score 73.8 | Mixture-of-Subspaces in Low-Rank Adaptation | wutaiqiang/moslora | 5 | Compare |
Papers archive 2025-07-28
4 shown of 4 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 384. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Mixture-of-Subspaces in Low-Rank Adaptation | 1 | 2 | 16 Jun 2024 | ran 4 of 6 samples (2 unverified; 6 pointer-only for licence) |
| CuMo: Scaling Multimodal LLM with Co-Upcycled Mixture-of-Experts | 1 | 1 | 9 May 2024 | ran 10 of 12 samples (2 unverified) |
| Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization | 1 | 1 | 5 Feb 2024 | ran 3 of 5 samples (2 unverified; 5 pointer-only for licence) |
| DreamLLM: Synergistic Multimodal Comprehension and Creation | 1 | 1 | 20 Sep 2023 | ran 3 of 3 samples (0 unverified; 2 pointer-only for licence) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- MMBench
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections