Datasets › SVAMP

SVAMP (Simple Variations on Arithmetic Math word Problems)

Introduced by Arkil Patel et al. in Are NLP Models really able to Solve Simple Math Word Problems?12 Mar 2021 archive 2025-07-28

A challenge set for elementary-level Math Word Problems (MWP). An MWP consists of a short Natural Language narrative that describes a state of the world and poses a question about some unknown quantities.

The examples in SVAMP test a model across different aspects of solving MWPs: 1) Is the model question sensitive? 2) Does the model have robust reasoning ability? 3) Is it invariant to structural alterations?

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Math Word Problem Solving SVAMP GPT-4 (Teaching-Inspired) Execution Accuracy 93.9 Teaching-Inspired Integrated Prompting Framework: A... sallytan13/teaching-inspired-prompting 26 Compare
Math Word Problem Solving SVAMP (1:N) ATHENA (roberta-large) Execution Accuracy 67.8 ATHENA: Mathematical Reasoning with Thought Expansion the-jb/athena-math 2 Compare

Papers archive 2025-07-28

16 shown of 16 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 362. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models 1 1 10 Oct 2024 ran 8 of 8 samples (0 unverified; 8 pointer-only for licence)
Achieving >97% on GSM8K: Deeply Understanding the Problems Makes LLMs Better Solvers for Math Word Problems 1 1 23 Apr 2024 not harvested
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning 1 3 23 Feb 2024 ran 9 of 11 samples (2 unverified; 11 pointer-only for licence)
OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset 1 1 15 Feb 2024 not harvested
Frugal LMs Trained to Invoke Symbolic Solvers Achieve Parameter-Efficient Arithmetic Reasoning 1 2 9 Dec 2023 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
ATHENA: Mathematical Reasoning with Thought Expansion 1 4 2 Nov 2023 not harvested
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning 1 1 5 Oct 2023 ran 2 of 2 samples (0 unverified)
Llama 2: Open Foundation and Fine-Tuned Chat Models 19 1 18 Jul 2023 ran 31 of 52 samples (21 unverified; 16 pointer-only for licence)
Math Word Problem Solving by Generating Linguistic Variants of Problem Statements 1 1 24 Jun 2023 not harvested
Does ChatGPT Comprehend the Place Value in Numbers When Solving Math Word Problems? 1 2 3 Jun 2023 not harvested
Learning Multi-Step Reasoning by Solving Arithmetic Tasks 1 1 2 Jun 2023 not harvested
Automatic Model Selection with Large Language Models for Reasoning 1 1 23 May 2023 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
Progressive-Hint Prompting Improves Reasoning in Large Language Models 1 1 19 Apr 2023 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
Large Language Models are Zero-Shot Reasoners 4 2 24 May 2022 ran 0 of 4 samples (4 unverified; 1 pointer-only for licence)
Learning to Reason Deductively: Math Word Problem Solving as Complex Relation Extraction 1 1 19 Mar 2022 ran 2 of 3 samples (1 unverified; 3 pointer-only for licence)
Are NLP Models really able to Solve Simple Math Word Problems? 3 4 12 Mar 2021 not harvested

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • SVAMP
  • SVAMP (1:N)

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections