Datasets › MMLU-Pro
MMLU-Pro
The MMLU-Pro dataset is an enhanced version of the Massive Multitask Language Understanding (MMLU) benchmark. It's designed to be more robust and challenging, aiming to rigorously benchmark large language models' capabilities in language comprehension and reasoning across diverse domains. Here are some key features of the MMLU-Pro dataset:
- Increased Complexity: It includes more reasoning-focused questions and expands the choice set from four to ten options, reducing the likelihood of random guessing and increasing the evaluation's complexity¹.
- Elimination of Trivial Questions: MMLU-Pro removes trivial and noisy questions found in the original MMLU, making it a more discriminative benchmark².
- Stability Under Varying Prompts: The dataset shows greater stability under varying prompts, with a decreased sensitivity of model scores to prompt variations².
- Better Performance with Reasoning: Models utilizing Chain of Thought (CoT) reasoning achieved better performance on MMLU-Pro compared to direct answering².
- Size and Scope: The dataset contains 12K complex questions across various disciplines¹⁴.
(1) TIGER-Lab/MMLU-Pro · Datasets at Hugging Face. https://huggingface.co/datasets/TIGER-Lab/MMLU-Pro. (2) MMLU-Pro: A More Robust and Challenging Multi-Task Language .... https://arxiv.org/abs/2406.01574. (3) MMLU-Pro: An Upgraded Version of the MMLU Dataset | LLM Explorer Blog. https://llm.extractum.io/static/blog/?id=mmlu-pro-benchmark. (4) TIGER-Lab Introduces MMLU-Pro Dataset for Comprehensive Benchmarking of .... https://www.marktechpost.com/2024/05/16/tiger-lab-introduces-mmlu-pro-dataset-for-comprehensive-benchmarking-of-large-language-models-capabilities-and-performance/. (5) undefined. https://doi.org/10.48550/arXiv.2406.01574.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| MMLU | MMLU-Pro | Orange-mini 0-shot MRR 99.19 | MyGO Multiplex CoT: A Method for Self-Reflection in... | data-dream-gdsp/Multiplex-CoT | 1 | Compare |
Papers archive 2025-07-28
1 shown of 1 paper with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 150. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| MyGO Multiplex CoT: A Method for Self-Reflection in Large Language Models via Double Chain of Thought Thinking | 1 | 1 | 20 Jan 2025 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
No modality tagged.
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- MMLU-Pro
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections