{"url":"/dataset/multi","name":"MULTI","full_name":null,"description_markdown":"MULTI-Benchmark is a cutting-edge benchmark for evaluating Multimodal Large Language Models (MLLMs). It is designed to test the understanding of complex tables and images, and reasoning with long context¹. Here are some key features of MULTI-Benchmark:\r\n\r\n- **Multimodal Inputs**: MULTI-Benchmark provides multimodal inputs and requires responses that are either precise or open-ended, reflecting real-life examination styles¹.\r\n- **Variety of Tasks**: It includes over 18,000 questions and challenges MLLMs with a variety of tasks, ranging from formula derivation to image detail analysis and cross-modality reasoning¹.\r\n- **MULTI-Elite and MULTI-Extend**: It introduces MULTI-Elite, a 500-question selected hard subset, and MULTI-Extend, with more than 4,500 external knowledge context pieces¹.\r\n- **Evaluation**: The evaluation indicates significant potential for MLLM advancement, with GPT-4V achieving a 63.7% accuracy rate on MULTI, in contrast to other MLLMs scoring between 28.5% and 55.3%¹.\r\n\r\n(1) ️ MULTI-Benchmark: Multimodal Understanding Leaderboard ... - GitHub. https://github.com/OpenDFM/MULTI-Benchmark.\r\n(2) ️ MULTI-Benchmark: Multimodal Understanding Leaderboard ... - GitHub. https://github.com/OpenDFM/MULTI-Benchmark.\r\n(3) ️ MULTI-Benchmark: Multimodal Understanding Leaderboard .... https://github.com/OpenDFM/MULTI-Benchmark/blob/main/README_zh.md.\r\n(4) MultiBench Dataset | Papers With Code. https://paperswithcode.com/dataset/multibench.\r\n(5) [2107.07502] MultiBench: Multiscale Benchmarks for Multimodal .... https://arxiv.org/abs/2107.07502.\r\n(6) MultiBench: Multiscale Benchmarks for Multimodal .... https://www.x-mol.com/paper/1416118156688826368?adv.\r\n(7) undefined. https://avatars.githubusercontent.com/u/139950066?v=4.\r\n(8) undefined. https://github.com/OpenDFM/MULTI-Benchmark/blob/main/README_zh.md?raw=true.\r\n(9) undefined. https://desktop.github.com.\r\n(10) undefined. https://docs.github.com/articles/about-issue-and-pull-request-templates.\r\n(11) undefined. https://github.com/OpenDFM/MULTI-Benchmark/raw/main/README_zh.md.\r\n(12) undefined. https://OpenDFM.github.io/MULTI-Benchmark/.\r\n(13) undefined. https://arxiv.org/abs/2402.03173/.\r\n(14) undefined. https://huggingface.co/datasets/OpenDFM/MULTI-Benchmark.","description_withheld":null,"homepage":"https://opendfm.github.io/MULTI-Benchmark","introduced_date":"2024-02-05","introduced_date_note":null,"introduced_by":{"paper":"/paper/multi-multimodal-understanding-leaderboard","title":"MULTI: Multimodal Understanding Leaderboard with Text and Images","first_author":"Zichen Zhu","url":null},"license":null,"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[],"languages":[{"name":"Chinese","url":"/datasets/language/chinese"}],"variants":["MULTI"],"data_loaders":[],"num_papers_in_archive":3,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[],"papers_with_a_benchmark_row":[],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}