{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmmu-a-massive-multi-discipline-multimodal","title":"MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI","arxiv_id":"2311.16502","date":"2023-11-27","proceeding":"CVPR 2024 1","authors":["Xiang Yue","Yuansheng Ni","Kai Zhang","Tianyu Zheng","Ruoqi Liu","Ge Zhang","Samuel Stevens","Dongfu Jiang","Weiming Ren","Yuxuan Sun","Cong Wei","Botao Yu","Ruibin Yuan","Renliang Sun","Ming Yin","Boyuan Zheng","Zhenzhu Yang","Yibo Liu","Wenhao Huang","Huan Sun","Yu Su","Wenhu Chen"],"abstract":"We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and textbooks, covering six core disciplines: Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly heterogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 14 open-source LMMs as well as the proprietary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence.","url_abs":"https://arxiv.org/abs/2311.16502v4","url_pdf":"https://arxiv.org/pdf/2311.16502v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mmmu-a-massive-multi-discipline-multimodal","repo_url":"https://github.com/MMMU-Benchmark/MMMU","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":null},{"paper_slug":"mmmu-a-massive-multi-discipline-multimodal","repo_url":"https://github.com/01-ai/yi","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"Apache-2.0"}},{"paper_slug":"mmmu-a-massive-multi-discipline-multimodal","repo_url":"https://github.com/eric-ai-lab/probmed","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"mmmu-a-massive-multi-discipline-multimodal","repo_url":"https://github.com/leloykun/mmfm-challenge","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"mmmu-a-massive-multi-discipline-multimodal","repo_url":"https://github.com/opendatalab/pm4bench","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[{"task_slug":"complex-query-answering","task_name":"Complex Query Answering"},{"task_slug":"logical-reasoning","task_name":"Logical Reasoning"},{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2311.16502","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2311.16502"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/01-ai/yi","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/opendatalab/pm4bench","reach":{"status":"ok","spdx":"Apache-2.0"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/eric-ai-lab/probmed","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/MMMU-Benchmark/MMMU","reach":null},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/leloykun/mmfm-challenge","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran_draft_wrong":4,"ran":1,"ran_honours":1,"unverified":8},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1},"listed":{"samples":11,"ran":3,"repositories":2}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"3f432bd90d038227","entry":"extract_subset_name","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"3f432bd90d038227"}},{"code_sha256_prefix":"42a46570620cd9fa","entry":"get_chunk","repo":"eric-ai-lab/probmed","repo_kind":"listed","path":"eval/inference/CheXagent/model_vqa_med.py","file_url":"https://github.com/eric-ai-lab/probmed/blob/HEAD/eval/inference/CheXagent/model_vqa_med.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"well_formed","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"42a46570620cd9fa"}},{"code_sha256_prefix":"829b3a29017bf6be","entry":"local_image_to_data_url","repo":"eric-ai-lab/probmed","repo_kind":"listed","path":"eval/inference/GPT-4V/gpt4v.py","file_url":"https://github.com/eric-ai-lab/probmed/blob/HEAD/eval/inference/GPT-4V/gpt4v.py","link_basis":"harvester_set","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"829b3a29017bf6be"}},{"code_sha256_prefix":"15b4072238264b35","entry":"mmmu_aggregate_results","repo":"MMMU-Benchmark/MMMU","repo_kind":"official","path":"mmmu-pro/evaluate.py","file_url":"https://github.com/MMMU-Benchmark/MMMU/blob/HEAD/mmmu-pro/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"15b4072238264b35"}},{"code_sha256_prefix":"9073174a2f7541e1","entry":"mmmu_process_results","repo":"MMMU-Benchmark/MMMU","repo_kind":"official","path":"mmmu-pro/evaluate.py","file_url":"https://github.com/MMMU-Benchmark/MMMU/blob/HEAD/mmmu-pro/evaluate.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"OUTPUT_MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"9073174a2f7541e1"}},{"code_sha256_prefix":"076c252c52cbb161","entry":"split_list","repo":"eric-ai-lab/probmed","repo_kind":"listed","path":"eval/inference/CheXagent/model_vqa_med.py","file_url":"https://github.com/eric-ai-lab/probmed/blob/HEAD/eval/inference/CheXagent/model_vqa_med.py","link_basis":"harvester_set","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":"deterministic","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"076c252c52cbb161"}},{"code_sha256_prefix":"c532e9d553008b9c","entry":"LLM_eval","repo":"leloykun/mmfm-challenge","repo_kind":"listed","path":"eval_only_mixtral.py","file_url":"https://github.com/leloykun/mmfm-challenge/blob/HEAD/eval_only_mixtral.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"c532e9d553008b9c"}},{"code_sha256_prefix":"30a511ce9cda0b48","entry":"calculate_ins_level_score","repo":"leloykun/mmfm-challenge","repo_kind":"listed","path":"eval_only_mixtral.py","file_url":"https://github.com/leloykun/mmfm-challenge/blob/HEAD/eval_only_mixtral.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"30a511ce9cda0b48"}},{"code_sha256_prefix":"d5c39813441ad22e","entry":"conv2sample","repo":"leloykun/mmfm-challenge","repo_kind":"listed","path":"eval_llava.py","file_url":"https://github.com/leloykun/mmfm-challenge/blob/HEAD/eval_llava.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"d5c39813441ad22e"}},{"code_sha256_prefix":"433752319ee12f66","entry":"creat_prompt","repo":"leloykun/mmfm-challenge","repo_kind":"listed","path":"eval_only_mixtral.py","file_url":"https://github.com/leloykun/mmfm-challenge/blob/HEAD/eval_only_mixtral.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"433752319ee12f66"}},{"code_sha256_prefix":"fc1c84f0076beee4","entry":"get_category_name","repo":"leloykun/mmfm-challenge","repo_kind":"listed","path":"eval_only.py","file_url":"https://github.com/leloykun/mmfm-challenge/blob/HEAD/eval_only.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"fc1c84f0076beee4"}},{"code_sha256_prefix":"dd2399dba3853e37","entry":"get_score_binary","repo":"eric-ai-lab/probmed","repo_kind":"listed","path":"eval/calculate_score.py","file_url":"https://github.com/eric-ai-lab/probmed/blob/HEAD/eval/calculate_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"dd2399dba3853e37"}},{"code_sha256_prefix":"e738c5a0630373d4","entry":"get_score_dict","repo":"eric-ai-lab/probmed","repo_kind":"listed","path":"eval/calculate_score.py","file_url":"https://github.com/eric-ai-lab/probmed/blob/HEAD/eval/calculate_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"e738c5a0630373d4"}},{"code_sha256_prefix":"eabb48a101284727","entry":"parse_response","repo":"eric-ai-lab/probmed","repo_kind":"listed","path":"eval/calculate_score.py","file_url":"https://github.com/eric-ai-lab/probmed/blob/HEAD/eval/calculate_score.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"eabb48a101284727"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}