Browse State-of-the-Art › MMLU
MMLU
156 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| MMLU-Pro (1 row) | Orange-mini | MyGO Multiplex CoT: A Method for Self-Reflection in Large Language... | code | — | Compare |
| mmlu (5-shots) (0 rows) | no rows in the archive | — | — | ||
| mmlu (chat CoT) (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 156 papers with code (340 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
20 Oct 2022 9 repositories listed Syntology ran 8 of 17 samples · 9 unverified · 2 pointer-only (licence)We find that instruction finetuning with the above aspects dramatically improves performance on a variety of model classes (PaLM, T5, U-PaLM), prompting setups (zero-shot, few-shot, CoT), and evaluation benchmarks…
-
18 Jun 2024 7 repositories listed Syntology ran 15 of 29 samples · 14 unverifiedWe introduce ChatGLM, an evolving family of large language models that we have been developing over time.
-
15 Jul 2024 6 repositories listedThis report introduces the Qwen2 series, the latest addition to our large language models and large multimodal models.
-
22 Feb 2024 4 repositories listed Syntology ran 10 of 39 samples · 29 unverified · 17 pointer-only (licence)The versatility of large language models (LLMs) led to the creation of diverse benchmarks that thoroughly test a variety of language models' abilities.
-
17 Jun 2024 3 repositories listed Syntology ran 22 of 22 samples · 0 unverifiedWe introduce DataComp for Language Models (DCLM), a testbed for controlled dataset experiments with the goal of improving language models.
-
6 Jun 2024 3 repositories listed Syntology ran 2 of 8 samples · 6 unverified · 1 pointer-only (licence)For example, we find that 57% of the analysed questions in the Virology subset contain errors.
-
30 Jan 2023 3 repositories listed Syntology ran 0 of 13 samples · 13 unverified · 13 pointer-only (licence)We introduce REPLUG, a retrieval-augmented language modeling framework that treats the language model (LM) as a black box and augments it with a tuneable retrieval model.
-
9 Feb 2025 2 repositories listedThis paper introduces the Large Memory Model (LM2), a decoder-only Transformer architecture enhanced with an auxiliary memory module that aims to address the limitations of standard Transformers in multi-step reasoning,…
-
22 Jun 2024 2 repositories listed Syntology ran 5 of 10 samples · 5 unverifiedDespite some recognition of redundancy in LLMs, the variability of redundancy across different architectures in transformers, such as MLP and Attention layers, is under-explored.
-
3 Jun 2024 2 repositories listed Syntology ran 9 of 12 samples · 3 unverifiedIn the age of large-scale language models, benchmarks like the Massive Multitask Language Understanding (MMLU) have been pivotal in pushing the boundaries of what AI can achieve in language comprehension and reasoning…
-
28 May 2024 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)The rapid advancements in Large Language Models (LLMs) have revolutionized various natural language processing tasks.
-
27 May 2024 2 repositories listed Syntology ran 7 of 9 samples · 2 unverifiedMost popular benchmarks for comparing LLMs rely on a limited set of prompt templates, which may not fully capture the LLMs' abilities and can affect the reproducibility of results on leaderboards.
-
25 Apr 2024 2 repositories listed Syntology ran 5 of 5 samples · 0 unverifiedWhile many contemporary large language models (LLMs) can process lengthy input, they still struggle to fully utilize information within the long context, known as the lost-in-the-middle challenge.
-
15 Feb 2024 2 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedTo create a benchmark, researchers must choose a dataset of forbidden prompts to which a victim model will respond, along with an evaluation method that scores the harmfulness of the victim model's responses.
-
19 Sep 2023 2 repositories listedLarge language models (LLMs) have demonstrated remarkable performance on a variety of natural language tasks based on just a few examples of natural language instructions, reducing the need for extensive feature…
-
18 Aug 2023 2 repositories listedIn this work, we propose a new safety evaluation benchmark RED-EVAL that carries out red-teaming.
-
27 May 2023 2 repositories listed Syntology ran 5 of 11 samples · 6 unverifiedRetrieval augmentation can aid language models (LMs) in knowledge-intensive tasks by supplying them with external information.
-
16 Mar 2023 2 repositories listed Syntology ran 0 of 4 samples · 4 unverifiedWe introduce Automatic Reasoning and Tool-use (ART), a framework that uses frozen LLMs to automatically generate intermediate reasoning steps as a program.
-
5 Aug 2022 2 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Retrieval augmented models are known to excel at knowledge intensive tasks without the need for as many parameters, but it is unclear whether they work in few-shot settings.
-
10 May 2022 2 repositories listed Syntology ran 0 of 16 samples · 16 unverifiedOur model also achieve strong results at in-context learning, outperforming 175B GPT-3 on zero-shot SuperGLUE and tripling the performance of T5-XXL on one-shot summarization.
-
29 Mar 2022 2 repositories listed Syntology ran 8 of 11 samples · 3 unverified · 4 pointer-only (licence)We investigate the optimal model size and number of tokens for training a transformer language model under a given compute budget.
-
15 Jul 2025 1 repository listedWe present Step-wise Policy for Rare-tool Knowledge (SPaRK), a novel reinforcement learning framework that teaches large language models to explore diverse tool usage patterns beyond conventional high-temperature…
-
8 Jul 2025 1 repository listedWe formulate the delta learning hypothesis to explain this phenomenon, positing that the relative quality delta between points suffices to drive learning via preference tuning--even when supervised finetuning on the…
-
8 Jul 2025 1 repository listedThe prevailing paradigm for scaling large language models (LLMs) involves monolithic, end-to-end training, a resource-intensive process that lacks flexibility.
-
7 Jul 2025 1 repository listedWe attribute this to "representational interference" in conventional models, where the embedding layer is burdened with learning both structural and semantic features.
-
7 Jul 2025 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)We present any4, a learned 4-bit weight quantization solution for large language models (LLMs) providing arbitrary numeric representations without requiring pre-processing of weights or activations.
-
31 May 2025 1 repository listedPsychology research has shown that humans are poor at estimating their performance on tasks, tending towards underconfidence on easy tasks and overconfidence on difficult tasks.
-
30 May 2025 1 repository listed Syntology ran 0 of 7 samples · 7 unverifiedWe thus introduce HELM, a family of HypErbolic Large Language Models, offering a geometric rethinking of the Transformer-based LLM that addresses the representational inflexibility, missing set of necessary operations,…
-
30 May 2025 1 repository listedThe performance of large language models (LLMs) continues to improve, as reflected in rising scores on standard benchmarks.
-
30 May 2025 1 repository listedIn the second stage, language descriptions are fed into a powerful reasoning LLM to solve complex video-language understanding tasks.
Syntology lines on 19 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections