{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/mmau-a-holistic-benchmark-of-agent","title":"MMAU: A Holistic Benchmark of Agent Capabilities Across Diverse Domains","arxiv_id":"2407.18961","date":"2024-07-18","proceeding":null,"authors":["Guoli Yin","Haoping Bai","Shuang Ma","Feng Nan","Yanchao Sun","Zhaoyang Xu","Shen Ma","Jiarui Lu","Xiang Kong","Aonan Zhang","Dian Ang Yap","Yizhe Zhang","Karsten Ahnert","Vik Kamath","Mathias Berglund","Dominic Walsh","Tobias Gindele","Juergen Wiest","Zhengfeng Lai","Xiaoming Wang","Jiulong Shan","Meng Cao","Ruoming Pang","ZiRui Wang"],"abstract":"Recent advances in large language models (LLMs) have increased the demand for comprehensive benchmarks to evaluate their capabilities as human-like agents. Existing benchmarks, while useful, often focus on specific application scenarios, emphasizing task completion but failing to dissect the underlying skills that drive these outcomes. This lack of granularity makes it difficult to deeply discern where failures stem from. Additionally, setting up these environments requires considerable effort, and issues of unreliability and reproducibility sometimes arise, especially in interactive tasks. To address these limitations, we introduce the Massive Multitask Agent Understanding (MMAU) benchmark, featuring comprehensive offline tasks that eliminate the need for complex environment setups. It evaluates models across five domains, including Tool-use, Directed Acyclic Graph (DAG) QA, Data Science and Machine Learning coding, Contest-level programming and Mathematics, and covers five essential capabilities: Understanding, Reasoning, Planning, Problem-solving, and Self-correction. With a total of 20 meticulously designed tasks encompassing over 3K distinct prompts, MMAU provides a comprehensive framework for evaluating the strengths and limitations of LLM agents. By testing 18 representative models on MMAU, we provide deep and insightful analyses. Ultimately, MMAU not only sheds light on the capabilities and limitations of LLM agents but also enhances the interpretability of their performance. Datasets and evaluation scripts of MMAU are released at https://github.com/apple/axlearn/tree/main/docs/research/mmau.","url_abs":"https://arxiv.org/abs/2407.18961v3","url_pdf":"https://arxiv.org/pdf/2407.18961v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"mmau-a-holistic-benchmark-of-agent","repo_url":"https://github.com/apple/axlearn","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"jax","reach":{"status":"ok","spdx":"Apache-2.0"}}],"tasks":[],"methods":[{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/2407.18961","atlas_url":"https://app.syntology.ai/?focus=2407.18961","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2407.18961"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-25T09:33:49+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/apple/axlearn","reach":{"status":"ok","spdx":"Apache-2.0"}}],"summary":{"ran":3},"by_repo_kind":{"official":{"samples":3,"ran":3,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"b6ee3869a2411cdb","entry":"no_remat","repo":"apple/axlearn","repo_kind":"official","path":"axlearn/common/base_layer.py","file_url":"https://github.com/apple/axlearn/blob/HEAD/axlearn/common/base_layer.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b6ee3869a2411cdb"}},{"code_sha256_prefix":"b37512092e0abd1b","entry":"nowrap","repo":"apple/axlearn","repo_kind":"official","path":"axlearn/common/module.py","file_url":"https://github.com/apple/axlearn/blob/HEAD/axlearn/common/module.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"b37512092e0abd1b"}},{"code_sha256_prefix":"c9415849ea23c425","entry":"round_features","repo":"apple/axlearn","repo_kind":"official","path":"axlearn/vision/mobilenets_blocks.py","file_url":"https://github.com/apple/axlearn/blob/HEAD/axlearn/vision/mobilenets_blocks.py","link_basis":"first_harvest_node","language":"python","status":"ran","verification_level":1,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"Apache-2.0","inline_ok":true,"mcp_get_code":{"code_sha256":"c9415849ea23c425"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}