Browse State-of-the-Art › MMLU › Papers, page 3
MMLU
Papers archive 2025-07-28
archive papers tagged: 340 · with a code link: 156 · where Syntology ran a sample: 78 (60 with a run with no instrument failure, 18 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (78 of 340 tagged: 60 with a run with no instrument failure, 18 where every run was a failure of Syntology's instrument)
Page 3 of 4: papers 201 to 300 of 340, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Efficient Model Development through Fine-tuning Transfer25 Mar 2025 0 repositories listed
-
Bias Evaluation and Mitigation in Retrieval-Augmented Medical Question-Answering Systems19 Mar 2025 0 repositories listed
-
SuperBPE: Space Travel for Language Models17 Mar 2025 0 repositories listed
-
LAG-MMLU: Benchmarking Frontier LLM Understanding in Latvian and Giriama14 Mar 2025 0 repositories listed
-
MMLU-ProX: A Multilingual Benchmark for Advanced Large Language Model Evaluation13 Mar 2025 0 repositories listed
-
Evaluating Mathematical Reasoning Across Large Language Models: A Fine-Grained Approach13 Mar 2025 0 repositories listed
-
Effectiveness of Zero-shot-CoT in Japanese Prompts9 Mar 2025 0 repositories listed
-
Leveraging Approximate Caching for Faster Retrieval-Augmented Generation7 Mar 2025 0 repositories listed
-
Correctness Coverage Evaluation for Medical Multiple-Choice Question Answering Based on the Enhanced Conformal Prediction Framework7 Mar 2025 0 repositories listed
-
Symbolic Mixture-of-Experts: Adaptive Skill-based Routing for Heterogeneous Reasoning7 Mar 2025 0 repositories listed
-
Universality of Layer-Level Entropy-Weighted Quantization Beyond Model Architecture and Size6 Mar 2025 0 repositories listed
-
KurTail : Kurtosis-based LLM Quantization3 Mar 2025 0 repositories listed
-
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering3 Mar 2025 0 repositories listed
-
PolyPrompt: Automating Knowledge Extraction from Multilingual Language Models with Dynamic Prompt Generation27 Feb 2025 0 repositories listed
-
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?26 Feb 2025 0 repositories listed
-
Efficient Federated Search for Retrieval-Augmented Generation26 Feb 2025 0 repositories listed
-
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks24 Feb 2025 0 repositories listed
-
Detecting Benchmark Contamination Through Watermarking24 Feb 2025 0 repositories listed
-
Distributional Scaling Laws for Emergent Capabilities24 Feb 2025 0 repositories listed
-
Evaluating Expert Contributions in a MoE LLM for Quiz-Based Tasks24 Feb 2025 0 repositories listed
-
Swallowing the Poison Pills: Insights from Vulnerability Disparity Among LLMs23 Feb 2025 0 repositories listed
-
Obliviate: Efficient Unmemorization for Protecting Intellectual Property in Large Language Models20 Feb 2025 0 repositories listed
-
Triangulating LLM Progress through Benchmarks, Games, and Cognitive Tests20 Feb 2025 0 repositories listed
-
None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks18 Feb 2025 0 repositories listed
-
Language Complexity Measurement as a Noisy Zero-Shot Proxy for Evaluating LLM Performance17 Feb 2025 0 repositories listed
-
Towards Fully Exploiting LLM Internal States to Enhance Knowledge Boundary Perception17 Feb 2025 0 repositories listed
-
Leveraging Uncertainty Estimation for Efficient LLM Routing16 Feb 2025 0 repositories listed
-
OctoTools: An Agentic Framework with Extensible Tools for Complex Reasoning16 Feb 2025 0 repositories listed
-
ORI: O Routing Intelligence14 Feb 2025 0 repositories listed
-
Cost-Saving LLM Cascades with Early Abstention13 Feb 2025 0 repositories listed
-
Selective Self-to-Supervised Fine-Tuning for Generalization in Large Language Models12 Feb 2025 0 repositories listed
-
Tokenization Standards for Linguistic Integrity: Turkish as a Benchmark10 Feb 2025 0 repositories listed
-
FRAMES: Boosting LLMs with A Four-Quadrant Multi-Stage Pretraining Strategy8 Feb 2025 0 repositories listed
-
Adapt-Pruner: Adaptive Structural Pruning for Efficient Small Language Model Training5 Feb 2025 0 repositories listed
-
Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?2 Feb 2025 0 repositories listed
-
IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding27 Jan 2025 0 repositories listed
-
HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI26 Jan 2025 0 repositories listed
-
Humanity's Last Exam24 Jan 2025 0 repositories listed
-
On the Reasoning Capacity of AI Models and How to Quantify It23 Jan 2025 0 repositories listed
-
Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs21 Jan 2025 0 repositories listed
-
DNA 1.0 Technical Report18 Jan 2025 0 repositories listed
-
Inference-Time-Compute: More Faithful? A Research Note14 Jan 2025 0 repositories listed
-
Unraveling Indirect In-Context Learning Using Influence Functions1 Jan 2025 0 repositories listed
-
Monty Hall and Optimized Conformal Prediction to Improve Decision-Making with LLMs31 Dec 2024 0 repositories listed
-
Setting Standards in Turkish NLP: TR-MMLU for Large Language Model Evaluation31 Dec 2024 0 repositories listed
-
SecBench: A Comprehensive Multi-Dimensional Benchmarking Dataset for LLMs in Cybersecurity30 Dec 2024 0 repositories listed
-
MixLLM: LLM Quantization with Global Mixed-precision between Output-features and Highly-efficient System Design19 Dec 2024 0 repositories listed
-
ChainRank-DPO: Chain Rank Direct Preference Optimization for LLM Rankers18 Dec 2024 0 repositories listed
-
Unveiling the Secret Recipe: A Guide For Supervised Fine-Tuning Small LLMs17 Dec 2024 0 repositories listed
-
Nanoscaling Floating-Point (NxFP): NanoMantissa, Adaptive Microexponents, and Code Recycling for Direct-Cast Compression of Large Language Models15 Dec 2024 0 repositories listed
-
LLM Distillation for Efficient Few-Shot Multiple Choice Question Answering13 Dec 2024 0 repositories listed
-
Global MMLU: Understanding and Addressing Cultural and Linguistic Biases in Multilingual Evaluation4 Dec 2024 0 repositories listed
-
Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset3 Dec 2024 0 repositories listed
-
The Vulnerability of Language Model Benchmarks: Do They Accurately Reflect True LLM Performance?2 Dec 2024 0 repositories listed
-
Improving Physics Reasoning in Large Language Models Using Mixture of Refinement Agents1 Dec 2024 0 repositories listed
-
Simple and Provable Scaling Laws for the Test-Time Compute of Large Language Models29 Nov 2024 0 repositories listed
-
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference27 Nov 2024 0 repositories listed
-
Predicting Emergent Capabilities by Finetuning25 Nov 2024 0 repositories listed
-
GenBFA: An Evolutionary Optimization Approach to Bit-Flip Attacks on LLMs21 Nov 2024 0 repositories listed
-
Learning from "Silly" Questions Improves Large Language Models, But Only Slightly21 Nov 2024 0 repositories listed
-
Real-time Adapting Routing (RAR): Improving Efficiency Through Continuous Learning in Software Powered by Layered Foundation Models14 Nov 2024 0 repositories listed
-
Reasoning Robustness of LLMs to Adversarial Typographical Errors8 Nov 2024 0 repositories listed
-
Watson: A Cognitive Observability Framework for the Reasoning of LLM-Powered Agents5 Nov 2024 0 repositories listed
-
Project MPG: towards a generalized performance benchmark for LLM capabilities28 Oct 2024 0 repositories listed
-
Adaptive Dense Reward: Understanding the Gap Between Action and Reward Space in Alignment23 Oct 2024 0 repositories listed
-
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning18 Oct 2024 0 repositories listed
-
G-Designer: Architecting Multi-agent Communication Topologies via Graph Neural Networks15 Oct 2024 0 repositories listed
-
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs15 Oct 2024 0 repositories listed
-
Towards Multilingual LLM Evaluation for European Languages11 Oct 2024 0 repositories listed
-
Upcycling Large Language Models into Mixture of Experts10 Oct 2024 0 repositories listed
-
Large Language Model Compression with Neural Architecture Search9 Oct 2024 0 repositories listed
-
Reasoning Paths Optimization: Learning to Reason and Explore From Diverse Paths7 Oct 2024 0 repositories listed
-
Continuous Approximations for Improving Quantization Aware Training of LLMs6 Oct 2024 0 repositories listed
-
BrainTransformers: SNN-LLM3 Oct 2024 0 repositories listed
-
Efficiently Deploying LLMs with Controlled Risk3 Oct 2024 0 repositories listed
-
DoPAMine: Domain-specific Pre-training Adaptation from seed-guided data Mining30 Sep 2024 0 repositories listed
-
Instance-adaptive Zero-shot Chain-of-Thought Prompting30 Sep 2024 0 repositories listed
-
Teuken-7B-Base & Teuken-7B-Instruct: Towards European LLMs30 Sep 2024 0 repositories listed
-
SSR: Alignment-Aware Modality Connector for Speech Language Models30 Sep 2024 0 repositories listed
-
Uncovering Latent Chain of Thought Vectors in Language Models21 Sep 2024 0 repositories listed
-
Bilingual Evaluation of Language Models on General Knowledge in University Entrance Exams with Minimal Contamination19 Sep 2024 0 repositories listed
-
GRIN: GRadient-INformed MoE18 Sep 2024 0 repositories listed
-
CPL: Critical Plan Step Learning Boosts LLM Generalization in Reasoning Tasks13 Sep 2024 0 repositories listed
-
Eir: Thai Medical Large Language Models13 Sep 2024 0 repositories listed
-
Selective Self-Rehearsal: A Fine-Tuning Approach to Improve Generalization in Large Language Models7 Sep 2024 0 repositories listed
-
Reasoning Beyond Bias: A Study on Counterfactual Prompting and Chain of Thought Reasoning16 Aug 2024 0 repositories listed
-
SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models16 Aug 2024 0 repositories listed
-
BOTS-LM: Training Large Language Models for Setswana5 Aug 2024 0 repositories listed
-
Networks of Networks: Complexity Class Principles Applied to Compound AI Systems Design23 Jul 2024 0 repositories listed
-
ALLaM: Large Language Models for Arabic and English22 Jul 2024 0 repositories listed
-
AgentInstruct: Toward Generative Teaching with Agentic Flows3 Jul 2024 0 repositories listed
-
Cost-Effective Proxy Reward Model Construction with On-Policy and Active Learning2 Jul 2024 0 repositories listed
-
Changing Answer Order Can Decrease MMLU Accuracy27 Jun 2024 0 repositories listed
-
Data Efficient Evaluation of Large Language Models and Text-to-Image Models via Adaptive Sampling21 Jun 2024 0 repositories listed
-
DEM: Distribution Edited Model for Training with Mixed Data Distributions21 Jun 2024 0 repositories listed
-
Optimised Grouped-Query Attention Mechanism for Transformers21 Jun 2024 0 repositories listed
-
Pistis-RAG: Enhancing Retrieval-Augmented Generation with Human Feedback21 Jun 2024 0 repositories listed
-
Understanding Finetuning for Factual Knowledge Extraction20 Jun 2024 0 repositories listed
-
Cultural Conditioning or Placebo? On the Effectiveness of Socio-Demographic Prompting17 Jun 2024 0 repositories listed
-
The Base-Rate Effect on LLM Benchmark Performance: Disambiguating Test-Taking Strategies from Benchmark Performance17 Jun 2024 0 repositories listed