Datasets › MRR-Benchmark

MRR-Benchmark (Multi-Modal Reading Benchmark)

Introduced by Jian Chen et al. in MMR: Evaluating Reading Ability of Large Multimodal Models26 Aug 2024 archive 2025-07-28

Multi-Modal Reading (MMR) Benchmark includes 550 annotated question-answer pairs across 11 distinct tasks involving texts, fonts, visual elements, bounding boxes, spatial relations, and grounding, with carefully designed evaluation metrics.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
MMR total MRR-Benchmark Claude 3.5 Sonnet Total Column Score 463 Claude 3.5 Sonnet Model Card Addendum — 14 Compare

Papers archive 2025-07-28

10 shown of 10 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 11. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Claude 3.5 Sonnet Model Card Addendum 0 1 24 Jun 2024 not harvested
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding 0 1 14 Jun 2024 not harvested
What matters when building vision-language models? 0 1 3 May 2024 not harvested
Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone 0 1 22 Apr 2024 not harvested
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks 2 2 21 Dec 2023 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Monkey: Image Resolution and Text Label Are Important Things for Large Multi-modal Models 1 1 11 Nov 2023 ran 1 of 1 samples (0 unverified)
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) 2 1 29 Sep 2023 not harvested
Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond 2 2 24 Aug 2023 ran 0 of 2 samples (2 unverified; 2 pointer-only for licence)
OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents 2 1 21 Jun 2023 ran 0 of 1 samples (1 unverified)
Visual Instruction Tuning 13 3 17 Apr 2023 ran 16 of 51 samples (35 unverified)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Apache License

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MRR-Benchmark

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections