{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmworld-towards-multi-discipline-multi","title":"MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos","arxiv_id":"2406.08407","date":"2024-06-12","proceeding":null,"authors":["Xuehai He","Weixi Feng","Kaizhi Zheng","Yujie Lu","Wanrong Zhu","Jiachen Li","Yue Fan","JianFeng Wang","Linjie Li","Zhengyuan Yang","Kevin Lin","William Yang Wang","Lijuan Wang","Xin Eric Wang"],"abstract":"Multimodal Language Language Models (MLLMs) demonstrate the emerging abilities of \"world models\" -- interpreting and reasoning about complex real-world dynamics. To assess these abilities, we posit videos are the ideal medium, as they encapsulate rich representations of real-world dynamics and causalities. To this end, we introduce MMWorld, a new benchmark for multi-discipline, multi-faceted multimodal video understanding. MMWorld distinguishes itself from previous video understanding benchmarks with two unique advantages: (1) multi-discipline, covering various disciplines that often require domain expertise for comprehensive understanding; (2) multi-faceted reasoning, including explanation, counterfactual thinking, future prediction, etc. MMWorld consists of a human-annotated dataset to evaluate MLLMs with questions about the whole videos and a synthetic dataset to analyze MLLMs within a single modality of perception. Together, MMWorld encompasses 1,910 videos across seven broad disciplines and 69 subdisciplines, complete with 6,627 question-answer pairs and associated captions. The evaluation includes 2 proprietary and 10 open-source MLLMs, which struggle on MMWorld (e.g., GPT-4V performs the best with only 52.3\\% accuracy), showing large room for improvement. Further ablation studies reveal other interesting findings such as models' different skill sets from humans. We hope MMWorld can serve as an essential step towards world model evaluation in videos.","url_abs":"https://arxiv.org/abs/2406.08407v3","url_pdf":"https://arxiv.org/pdf/2406.08407v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mmworld-towards-multi-discipline-multi","repo_url":"https://github.com/eric-ai-lab/mmworld","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"future-prediction","task_name":"Future prediction"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":null,"task_name":"counterfactual"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2406.08407","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2406.08407"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/eric-ai-lab/mmworld","reach":{"status":"ok","spdx":"MIT"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"00976758a2225680","entry":"answer_post_processing","repo":"eric-ai-lab/mmworld","repo_kind":"official","path":"evaluation/main_utils.py","file_url":"https://github.com/eric-ai-lab/mmworld/blob/HEAD/evaluation/main_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"00976758a2225680"}},{"code_sha256_prefix":"acca92400119d7a2","entry":"calculate_video_length","repo":"eric-ai-lab/mmworld","repo_kind":"official","path":"evaluation/main_utils.py","file_url":"https://github.com/eric-ai-lab/mmworld/blob/HEAD/evaluation/main_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"acca92400119d7a2"}},{"code_sha256_prefix":"0eb451ada078cf49","entry":"format_time","repo":"eric-ai-lab/mmworld","repo_kind":"official","path":"evaluation/main_utils.py","file_url":"https://github.com/eric-ai-lab/mmworld/blob/HEAD/evaluation/main_utils.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"0eb451ada078cf49"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}