Datasets › GSM8K

GSM8K

Introduced by Karl Cobbe et al. in Training Verifiers to Solve Math Word Problems27 Oct 2021 archive 2025-07-28

GSM8K is a dataset of 8.5K high quality linguistically diverse grade school math word problems created by human problem writers. The dataset is segmented into 7.5K training problems and 1K test problems. These problems take between 2 and 8 steps to solve, and solutions primarily involve performing a sequence of elementary calculations using basic arithmetic operations (+ − ×÷) to reach the final answer. A bright middle school student should be able to solve every problem. It can be used for multi-step mathematical reasoning.

Image source: https://arxiv.org/pdf/2110.14168v1.pdf

Benchmarks archive 2025-07-28

All 7 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Arithmetic Reasoning GSM8K Claude 3.5 Sonnet (HPT) Accuracy 97.72 Hierarchical Prompting Taxonomy: A Universal Evaluation... devichand579/HPT 164 Compare
GSM8K GSM8K Xolver Accuracy 98.1 Xolver: Multi-Agent Reasoning with Holistic Experience... kagnlp/Xolver 3 Compare
GSM8K gsm8k (5-shots) no rows — — 0 Compare
Mathematical Reasoning GSM8K no rows — — 0 Compare
Text Generation GSM8k (5-shot) no rows — — 0 Compare
Text Generation GSM8k TR no rows — — 0 Compare
Text Generation GSM8k TR v0.2 no rows — — 0 Compare

Papers archive 2025-07-28

30 shown of 58 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 1,881. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Xolver: Multi-Agent Reasoning with Holistic Experience Learning Just Like an Olympiad Team 1 1 17 Jun 2025 not harvested
CAPO: Cost-Aware Prompt Optimization 2 3 22 Apr 2025 not harvested
MyGO Multiplex CoT: A Method for Self-Reflection in Large Language Models via Double Chain of Thought Thinking 1 1 20 Jan 2025 not harvested
Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models 1 1 10 Oct 2024 ran 8 of 8 samples (0 unverified; 8 pointer-only for licence)
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data 1 4 2 Oct 2024 ran 0 of 15 samples (15 unverified)
Qwen2 Technical Report 6 1 15 Jul 2024 not harvested
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs 1 1 26 Jun 2024 ran 7 of 12 samples (5 unverified; 12 pointer-only for licence)
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles 1 1 18 Jun 2024 not harvested
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 1 8 18 Jun 2024 ran 10 of 12 samples (2 unverified)
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling 1 1 18 Jun 2024 ran 4 of 7 samples (3 unverified)
Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems 1 1 23 Apr 2024 not harvested
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing 1 2 18 Apr 2024 ran 13 of 18 samples (5 unverified; 18 pointer-only for licence)
Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM 1 1 12 Mar 2024 not harvested
The Claude 3 Model Family: Opus, Sonnet, Haiku 0 3 4 Mar 2024 not harvested
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning 1 4 23 Feb 2024 ran 9 of 11 samples (2 unverified; 11 pointer-only for licence)
Orca-Math: Unlocking the potential of SLMs in Grade School Math 0 1 16 Feb 2024 not harvested
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset 1 12 15 Feb 2024 not harvested
The Unreasonable Effectiveness of Eccentric Automatic Prompts 0 3 9 Feb 2024 not harvested
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 5 1 5 Feb 2024 ran 8 of 24 samples (16 unverified)
Parameter-Efficient Sparsity Crafting from Dense to Mixture-of-Experts for Instruction Tuning on General Tasks 2 2 5 Jan 2024 not harvested
Gemini: A Family of Highly Capable Multimodal Models 1 1 19 Dec 2023 not harvested
TinyGSM: achieving >80% on GSM8k with small language models 0 2 14 Dec 2023 not harvested
Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations 3 2 14 Dec 2023 not harvested
Fewer is More: Boosting LLM Reasoning with Reinforced Context Pruning 0 1 14 Dec 2023 not harvested
Frugal LMs Trained to Invoke Symbolic Solvers Achieve Parameter-Efficient Arithmetic Reasoning 1 1 9 Dec 2023 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Orca 2: Teaching Small Language Models How to Reason 0 2 18 Nov 2023 not harvested
OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning 1 3 16 Nov 2023 ran 3 of 3 samples (0 unverified; 3 pointer-only for licence)
The ART of LLM Refinement: Ask, Refine, and Trust 0 1 14 Nov 2023 not harvested
Llemma: An Open Language Model For Mathematics 4 2 16 Oct 2023 ran 6 of 8 samples (2 unverified)
KwaiYiiMath: Technical Report 0 1 11 Oct 2023 not harvested

The full list of 58 is in the JSON twin.

Dataset loaders archive 2025-07-28

22 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • GSM8K
  • GSM8k (5-shot)
  • GSM8k TR v0.2
  • GSM8k TR
  • gsm8k (5-shots)

5 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections