Datasets › MML

MML (Massive Multitask Language Understanding)

Introduced by Dan Hendrycks et al. in Measuring Massive Multitask Language Understanding7 Sep 2020 archive 2025-07-28

MMLU (Massive Multitask Language Understanding) is a new benchmark designed to measure knowledge acquired during pretraining by evaluating models exclusively in zero-shot and few-shot settings. This makes the benchmark more challenging and more similar to how we evaluate humans. The benchmark covers 57 subjects across STEM, the humanities, the social sciences, and more. It ranges in difficulty from an elementary level to an advanced professional level, and it tests both world knowledge and problem solving ability. Subjects range from traditional areas, such as mathematics and history, to more specialized areas like law and ethics. The granularity and breadth of the subjects makes the benchmark ideal for identifying a model’s blind spots.

Image source: https://arxiv.org/pdf/2009.03300v3.pdf

Benchmarks archive 2025-07-28

All 29 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Multi-task Language Understanding MML GPT-4 o1(300b) Average (%) 87 GPT-4o as the Gold Standard: A Scalable and General... — 44 Compare
Multiple Choice Question Answering (MCQA) MMLU (College Biology) Med-PaLM 2 (ER) Accuracy 95.8 Towards Expert-Level Medical Question Answering with... m42-health/med42 8 Compare
Multiple Choice Question Answering (MCQA) MMLU (Medical Genetics) Med-PaLM 2 (ER) Accuracy 92 Towards Expert-Level Medical Question Answering with... m42-health/med42 8 Compare
Multiple Choice Question Answering (MCQA) MMLU (Professional medicine) Med-PaLM 2 (5-shot) Accuracy 95.2 Towards Expert-Level Medical Question Answering with... m42-health/med42 6 Compare
Multiple Choice Question Answering (MCQA) MMLU (Elementary Mathematics) Chinchilla (few-shot, k=5) Accuracy 41.5 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (High School Biology) Chinchilla (few-shot, k=5) Accuracy 80.3 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (College Chemistry) Chinchilla (few-shot, k=5) Accuracy 51 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (High School Mathematics) GAL 120B (zero-shot) Accuracy 32.6 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (Electrical Engineer) GAL 120B (zero-shot) Accuracy 62.8 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (College Physics) Chinchilla (few-shot, k=5) Accuracy 46.1 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (Formal Logic) Gopher (few-shot, k=5) Accuracy 35.7 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (High School Statistics) Chinchilla (few-shot, k=5) Accuracy 58.8 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (Abstract Algebra) GAL 30B (zero-shot) Accuracy 33.3 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (Econometrics) Gopher (few-shot, k=5) Accuracy 43 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (High School Computer Science) GAL 120B (zero-shot) Accuracy 70 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (College Mathematics) GAL 120B (zero-shot) Accuracy 43 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (Astronomy) Chinchilla (few-shot, k=5) Accuracy 73.0 Galactica: A Large Language Model for Science paperswithcode/galai 5 Compare
Multiple Choice Question Answering (MCQA) MMLU (High School Chemistry) Chinchilla (few-shot, k=5) Accuracy 58.1 Galactica: A Large Language Model for Science paperswithcode/galai 4 Compare
Multiple Choice Question Answering (MCQA) MMLU (College Computer Science) Chinchilla (few-shot, k=5) Accuracy 51.0 Galactica: A Large Language Model for Science paperswithcode/galai 4 Compare
Multiple Choice Question Answering (MCQA) MMLU (High School Physics) Chinchilla (few-shot, k=5) Accuracy 36.4 Galactica: A Large Language Model for Science paperswithcode/galai 4 Compare
Multiple Choice Question Answering (MCQA) MMLU (Machine Learning) Chinchilla (few-shot, k=5) Accuracy 41.1 Galactica: A Large Language Model for Science paperswithcode/galai 4 Compare
Multiple Choice Question Answering (MCQA) MMLU (Clinical Knowledge) Med-PaLM 2 (ER) Accuracy 88.7 Towards Expert-Level Medical Question Answering with... m42-health/med42 3 Compare
Multiple Choice Question Answering (MCQA) MMLU (Anatomy) Med-PaLM 2 (ER) Accuracy 84.4 Towards Expert-Level Medical Question Answering with... m42-health/med42 3 Compare
Multiple Choice Question Answering (MCQA) MMLU (College Medicine) Med-PaLM (ER) Accuracy 83.2 Towards Expert-Level Medical Question Answering with... m42-health/med42 3 Compare
Multi-task Language Understanding MMLU (5-Shot) Sakalti/ultiima-78B MMLU (5-shot) 89.2 MERGE: Fast Private Text Generation liangzid/MERGE 1 Compare
Question Answering MML qwen-LLM 7B Accuracy 71.8 SHAKTI: A 2.5 Billion Parameter Small Language Model... — 1 Compare
Text Generation MMLU (5-Shot) no rows — — 0 Compare
Text Generation MMLU TR no rows — — 0 Compare
Text Generation MMLU TR v0.2 no rows — — 0 Compare

Papers archive 2025-07-28

25 shown of 25 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,922. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Llama 3 Meets MoE: Efficient Upcycling 1 2 13 Dec 2024 not harvested
SHAKTI: A 2.5 Billion Parameter Small Language Model Optimized for Edge AI and Low-Resource Environments 0 1 15 Oct 2024 not harvested
GPT-4o as the Gold Standard: A Scalable and General Purpose Approach to Filter Language Model Pretraining Data 0 1 3 Oct 2024 not harvested
The Llama 3 Herd of Models 5 2 31 Jul 2024 ran 2 of 9 samples (7 unverified)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM 1 1 12 Mar 2024 not harvested
The Claude 3 Model Family: Opus, Sonnet, Haiku 0 2 4 Mar 2024 not harvested
Mixtral of Experts 6 2 8 Jan 2024 ran 5 of 5 samples (0 unverified)
The Falcon Series of Open Language Models 0 2 28 Nov 2023 not harvested
MiLe Loss: a New Loss for Mitigating the Bias of Learning Difficulties in Generative Language Models 1 1 30 Oct 2023 not harvested
Mistral 7B 6 1 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
Textbooks Are All You Need II: phi-1.5 technical report 1 1 11 Sep 2023 not harvested
BioMedGPT: Open Multimodal Generative Pre-trained Transformer for BioMedicine 1 1 18 Aug 2023 not harvested
Llama 2: Open Foundation and Fine-Tuned Chat Models 19 5 18 Jul 2023 ran 31 of 52 samples (21 unverified; 16 pointer-only for licence)
MERGE: Fast Private Text Generation 1 1 25 May 2023 not harvested
Towards Expert-Level Medical Question Answering with Large Language Models 1 18 16 May 2023 not harvested
BloombergGPT: A Large Language Model for Finance 2 3 30 Mar 2023 not harvested
GPT-4 Technical Report 11 1 15 Mar 2023 ran 2 of 5 samples (3 unverified; 1 pointer-only for licence)
LLaMA: Open and Efficient Foundation Language Models 57 3 27 Feb 2023 ran 26 of 58 samples (32 unverified; 4 pointer-only for licence)
Galactica: A Large Language Model for Science 1 92 16 Nov 2022 ran 0 of 2 samples (2 unverified)
Scaling Instruction-Finetuned Language Models 9 9 20 Oct 2022 ran 8 of 17 samples (9 unverified; 2 pointer-only for licence)
GLM-130B: An Open Bilingual Pre-trained Model 9 1 5 Oct 2022 ran 5 of 21 samples (16 unverified)
Atlas: Few-shot Learning with Retrieval Augmented Language Models 2 1 5 Aug 2022 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
UL2: Unifying Language Learning Paradigms 2 1 10 May 2022 ran 0 of 16 samples (16 unverified)
GPT-NeoX-20B: An Open-Source Autoregressive Language Model 11 1 14 Apr 2022 ran 0 of 3 samples (3 unverified)
Training Compute-Optimal Large Language Models 2 1 29 Mar 2022 ran 8 of 11 samples (3 unverified; 4 pointer-only for licence)

Dataset loaders archive 2025-07-28

11 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Custom

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • mmlu
  • MMLU TR v0.2
  • MMLU TR
  • MMLU (5-Shot)
  • MMLU (College Medicine)
  • MMLU (Professional medicine)
  • MMLU (Anatomy)
  • MMLU (Clinical Knowledge)
  • MMLU (Medical Genetics)
  • MMLU (Mathematics)
  • MMLU (Machine Learning)
  • MMLU (High School Statistics)
  • MMLU (High School Physics)
  • MMLU (High School Mathematics)
  • MMLU (High School Computer Science)
  • MMLU (High School Chemistry)
  • MMLU (High School Biology)
  • MMLU (Formal Logic)
  • MMLU (Elementary Mathematics)
  • MMLU (Electrical Engineer)
  • MMLU (Econometrics)
  • MMLU (College Physics)
  • MMLU (College Mathematics)
  • MMLU (College Computer Science)
  • MMLU (College Chemistry)
  • MMLU (College Biology)
  • MMLU (Astronomy)
  • MMLU (Abstract Algebra)
  • MML

29 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections