{"url":"/dataset/mmlu-pro","name":"MMLU-Pro","full_name":null,"description_markdown":"The MMLU-Pro dataset is an enhanced version of the Massive Multitask Language Understanding (MMLU) benchmark. It's designed to be more robust and challenging, aiming to rigorously benchmark large language models' capabilities in language comprehension and reasoning across diverse domains. Here are some key features of the MMLU-Pro dataset:\r\n\r\n- **Increased Complexity**: It includes more reasoning-focused questions and expands the choice set from four to ten options, reducing the likelihood of random guessing and increasing the evaluation's complexity¹.\r\n- **Elimination of Trivial Questions**: MMLU-Pro removes trivial and noisy questions found in the original MMLU, making it a more discriminative benchmark².\r\n- **Stability Under Varying Prompts**: The dataset shows greater stability under varying prompts, with a decreased sensitivity of model scores to prompt variations².\r\n- **Better Performance with Reasoning**: Models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering².\r\n- **Size and Scope**: The dataset contains 12K complex questions across various disciplines¹⁴.\r\n\r\n(1) TIGER-Lab/MMLU-Pro · Datasets at Hugging Face. https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro.\r\n(2) MMLU-Pro: A More Robust and Challenging Multi-Task Language .... https://arxiv.org/abs/2406.01574.\r\n(3) MMLU-Pro: An Upgraded Version of the MMLU Dataset | LLM Explorer Blog. https://llm.extractum.io/static/blog/?id=mmlu-pro-benchmark.\r\n(4) TIGER-Lab Introduces MMLU-Pro Dataset for Comprehensive Benchmarking of .... https://www.marktechpost.com/2024/05/16/tiger-lab-introduces-mmlu-pro-dataset-for-comprehensive-benchmarking-of-large-language-models-capabilities-and-performance/.\r\n(5) undefined. https://doi.org/10.48550/arXiv.2406.01574.","description_withheld":null,"homepage":"https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro","introduced_date":"2024-06-03","introduced_date_note":null,"introduced_by":{"paper":"/paper/mmlu-pro-a-more-robust-and-challenging-multi","title":"MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark","first_author":"YuBo Wang","url":null},"license":null,"modalities":[],"tasks":[{"name":"MMLU","url":"/task/mmlu","datasets_with_task":"/datasets/task/mmlu"}],"languages":[],"variants":["MMLU-Pro"],"data_loaders":[],"num_papers_in_archive":150,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/mmlu-on-mmlu-pro","task":"MMLU","dataset_variant":"MMLU-Pro","rows":1,"metrics":["0-shot MRR"],"first_row_in_archive_order":{"model":"Orange-mini","paper":"/paper/mygo-multiplex-cot-a-method-for-self","metrics":{"0-shot MRR":"99.19"},"code_links":[{"title":"data-dream-gdsp/Multiplex-CoT","url":"https://github.com/data-dream-gdsp/Multiplex-CoT"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/mygo-multiplex-cot-a-method-for-self","title":"MyGO Multiplex CoT: A Method for Self-Reflection in Large Language Models via Double Chain of Thought Thinking","date":"2025-01-20","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":0,"samples_harvested":0,"samples_ran":0,"samples_unverified":0,"pointer_only_for_licence":0,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}