Browse State-of-the-Art › MMLU › Papers, page 4
MMLU
Papers archive 2025-07-28
archive papers tagged: 340 · with a code link: 156 · where Syntology ran a sample: 78 (60 with a run with no instrument failure, 18 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (78 of 340 tagged: 60 with a run with no instrument failure, 18 where every run was a failure of Syntology's instrument)
Page 4 of 4: papers 301 to 340 of 340, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Reactor Mk.1 performances: MMLU, HumanEval and BBH test results15 Jun 2024 0 repositories listed
-
MMLU-SR: A Benchmark for Stress-Testing Reasoning Capability of Large Language Models15 Jun 2024 0 repositories listed
-
GEB-1.3B: Open Lightweight Large Language Model14 Jun 2024 0 repositories listed
-
Quantifying Variance in Evaluation Benchmarks14 Jun 2024 0 repositories listed
-
Does your data spark joy? Performance gains from domain upsampling at the end of training5 Jun 2024 0 repositories listed
-
3 Jun 2024 0 repositories listed
-
Spanish and LLM Benchmarks: is MMLU Lost in Translation?28 May 2024 0 repositories listed
-
GECKO: Generative Language Model for English, Code and Korean24 May 2024 0 repositories listed
-
An Assessment of Model-On-Model Deception10 May 2024 0 repositories listed
-
SUTRA: Scalable Multilingual Language Model Architecture7 May 2024 0 repositories listed
-
Octopus v4: Graph of language models30 Apr 2024 0 repositories listed
-
22 Apr 2024 0 repositories listed
-
Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models18 Apr 2024 0 repositories listed
-
LLaMA-Excitor: General Instruction Tuning via Indirect Feature Interaction1 Apr 2024 0 repositories listed
-
NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning30 Mar 2024 0 repositories listed
-
Few-Shot Recalibration of Language Models27 Mar 2024 0 repositories listed
-
CodingTeachLLM: Empowering LLM's Coding Ability via AST Prior Knowledge13 Mar 2024 0 repositories listed
-
4 Mar 2024 0 repositories listed
-
KorMedMCQA: Multi-Choice Question Answering Benchmark for Korean Healthcare Professional Licensing Examinations3 Mar 2024 0 repositories listed
-
OpenMedLM: Prompt engineering can out-perform fine-tuning in medical question-answering with open-source large language models29 Feb 2024 0 repositories listed
-
Do Large Language Models Mirror Cognitive Language Processing?28 Feb 2024 0 repositories listed
-
ARL2: Aligning Retrievers for Black-box Large Language Models via Self-guided Adaptive Relevance Labeling21 Feb 2024 0 repositories listed
-
Automating Dataset Updates Towards Reliable and Timely Evaluation of Large Language Models19 Feb 2024 0 repositories listed
-
Towards Uncertainty-Aware Language Agent25 Jan 2024 0 repositories listed
-
LLaMA Beyond English: An Empirical Study on Language Capability Transfer2 Jan 2024 0 repositories listed
-
Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities22 Dec 2023 0 repositories listed
-
YAYI 2: Multilingual Open-Source Large Language Models22 Dec 2023 0 repositories listed
-
AcademicGPT: Empowering Academic Research21 Nov 2023 0 repositories listed
-
Investigating Data Contamination in Modern Benchmarks for Large Language Models16 Nov 2023 0 repositories listed
-
ConceptPsy:A Benchmark Suite with Conceptual Comprehensiveness in Psychology16 Nov 2023 0 repositories listed
-
The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback31 Oct 2023 0 repositories listed
-
TeacherLM: Teaching to Fish Rather Than Giving the Fish, Language Modeling Likewise29 Oct 2023 0 repositories listed
-
Evaluation of large language models using an Indian language LGBTI+ lexicon26 Oct 2023 0 repositories listed
-
Irreducible Curriculum for Language Model Pretraining23 Oct 2023 0 repositories listed
-
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models9 Oct 2023 0 repositories listed
-
Pruning Large Language Models via Accuracy Predictor18 Sep 2023 0 repositories listed
-
The Poison of Alignment25 Aug 2023 0 repositories listed
-
Let's Do a Thought Experiment: Using Counterfactuals to Improve Moral Reasoning25 Jun 2023 0 repositories listed
-
Measuring Progress on Scalable Oversight for Large Language Models4 Nov 2022 0 repositories listed
-
20 Oct 2022 0 repositories listed