Datasets › MATH

MATH

Introduced by Dan Hendrycks et al. in Measuring Mathematical Problem Solving With the MATH Dataset5 Mar 2021 archive 2025-07-28

MATH is a new dataset of 12,500 challenging competition mathematics problems. Each problem in MATH has a full step-by-step solution which can be used to teach models to generate answer derivations and explanations.

Source: Hendrycks et al.

Image source: Hendrycks et al.

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Math Word Problem Solving MATH Gemini 2.0 Flash Experimental Accuracy 89.7 — — 135 Compare
Math Word Problem Solving MATH minival Process Supervision (GPT-4) Accuracy 78.2 Let's Verify Step by Step openai/prm800k +2 1 Compare

Papers archive 2025-07-28

30 shown of 37 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,330. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data 1 4 2 Oct 2024 ran 0 of 15 samples (15 unverified)
Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement 0 6 18 Sep 2024 not harvested
Qwen2 Technical Report 6 1 15 Jul 2024 not harvested
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs 1 1 26 Jun 2024 ran 7 of 12 samples (5 unverified; 12 pointer-only for licence)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 1 8 18 Jun 2024 ran 10 of 12 samples (2 unverified)
AlphaMath Almost Zero: Process Supervision without Process 1 1 6 May 2024 not harvested
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing 1 1 18 Apr 2024 ran 13 of 18 samples (5 unverified; 18 pointer-only for licence)
MACM: Utilizing a Multi-Agent System for Condition Mining in Solving Complex Mathematical Problems 1 1 6 Apr 2024 ran 0 of 1 samples (1 unverified; 1 pointer-only for licence)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM 1 1 12 Mar 2024 not harvested
Key-Point-Driven Data Synthesis with its Enhancement on Mathematical Reasoning 0 4 4 Mar 2024 not harvested
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning 1 4 23 Feb 2024 ran 9 of 11 samples (2 unverified; 11 pointer-only for licence)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset 1 12 15 Feb 2024 not harvested
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 5 2 5 Feb 2024 ran 8 of 24 samples (16 unverified)
Augmenting Math Word Problems via Iterative Question Composing 1 1 17 Jan 2024 ran 4 of 4 samples (0 unverified; 4 pointer-only for licence)
Mixtral of Experts 6 2 8 Jan 2024 ran 5 of 5 samples (0 unverified)
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks 2 2 5 Jan 2024 not harvested
Gemini: A Family of Highly Capable Multimodal Models 1 2 19 Dec 2023 not harvested
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations 3 3 14 Dec 2023 not harvested
Mistral 7B 6 1 10 Oct 2023 ran 9 of 11 samples (2 unverified; 1 pointer-only for licence)
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning 1 3 9 Oct 2023 not harvested
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning 1 6 5 Oct 2023 ran 2 of 2 samples (0 unverified)
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving 1 8 29 Sep 2023 ran 5 of 12 samples (7 unverified)
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models 1 3 21 Sep 2023 ran 15 of 22 samples (7 unverified)
OpenChat: Advancing Open-source Language Models with Mixed-Quality Data 1 2 20 Sep 2023 ran 0 of 1 samples (1 unverified)
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct 1 4 18 Aug 2023 ran 10 of 16 samples (6 unverified; 16 pointer-only for licence)
Solving Challenging Math Word Problems Using GPT-4 Code Interpreter with Code-based Self-Verification 1 5 15 Aug 2023 not harvested
Cumulative Reasoning with Large Language Models 1 2 8 Aug 2023 ran 5 of 6 samples (1 unverified; 6 pointer-only for licence)
Skills-in-Context Prompting: Unlocking Compositionality in Large Language Models 0 1 1 Aug 2023 not harvested
Let's Verify Step by Step 3 1 31 May 2023 ran 3 of 3 samples (0 unverified)
PaLM 2 Technical Report 1 2 17 May 2023 not harvested

The full list of 37 is in the JSON twin.

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • MATH
  • MATH minival

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections