Browse State-of-the-Art › Math Word Problem Solving
Math Word Problem Solving
80 papers with code · 13 benchmarks · 23 datasets archive 2025-07-28
A math word problem is a mathematical exercise (such as in a textbook, worksheet, or exam) where significant background information on the problem is presented in ordinary language rather than in mathematical notation. As most word problems involve a narrative of some sort, they are sometimes referred to as story problems and may vary in the amount of technical language used.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
13 leaderboard tables shown for this task, 13 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 13 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
23 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 80 papers with code (107 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
27 Feb 2023 57 repositories listed Syntology ran 26 of 58 samples · 32 unverified · 4 pointer-only (licence)We introduce LLaMA, a collection of foundation language models ranging from 7B to 65B parameters.
-
18 Jul 2023 19 repositories listed Syntology ran 31 of 52 samples · 21 unverified · 16 pointer-only (licence)In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters.
-
5 Jun 2020 14 repositories listed Syntology ran 4 of 13 samples · 9 unverified · 3 pointer-only (licence)Recent progress in pre-trained neural language models has significantly improved the performance of many natural language processing (NLP) tasks.
-
2 Apr 2019 7 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedThe structured nature of the mathematics domain, covering arithmetic, algebra, probability and calculus, enables the construction of training and test splits designed to clearly illuminate the capabilities and…
-
15 Jul 2024 6 repositories listedThis report introduces the Qwen2 series, the latest addition to our large language models and large multimodal models.
-
8 Jan 2024 6 repositories listed Syntology ran 5 of 5 samples · 0 unverifiedIn particular, Mixtral vastly outperforms Llama 2 70B on mathematics, code generation, and multilingual benchmarks.
-
10 Oct 2023 6 repositories listed Syntology ran 9 of 11 samples · 2 unverified · 1 pointer-only (licence)We introduce Mistral 7B v0.
-
5 Feb 2024 5 repositories listed Syntology ran 8 of 24 samples · 16 unverifiedMathematical reasoning poses a significant challenge for language models due to its complex and structured nature.
-
5 Mar 2021 5 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics.
-
24 May 2022 4 repositories listed Syntology ran 0 of 4 samples · 4 unverified · 1 pointer-only (licence)Pretrained large language models (LLMs) are widely used in many sub-fields of natural language processing (NLP) and generally known as excellent few-shot learners with task-specific exemplars.
-
14 Dec 2023 3 repositories listedIn this paper, we present an innovative process-oriented math process reward model called \textbf{Math-Shepherd}, which assigns a reward score to each step of math problem solutions.
-
31 May 2023 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedWe conduct our own investigation, finding that process supervision significantly outperforms outcome supervision for training models to solve problems from the challenging MATH dataset.
-
22 Mar 2023 3 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedWe contend that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models.
-
18 Nov 2022 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedMuch of this success can be attributed to prompting methods such as "chain-of-thought'', which employ LLMs for both understanding the problem description by decomposing it into steps, as well as solving each step of the…
-
12 Mar 2021 3 repositories listedSince existing solvers achieve high performance on the benchmark datasets for elementary level MWPs containing one-unknown arithmetic word problems, such problems are often considered "solved" with the bulk of research…
-
5 Jan 2024 2 repositories listedUsing PESC during instruction tuning, our best sparse model outperforms other sparse and dense models and exhibits superior general capabilities compared to GPT-3.
-
2 Jun 2023 2 repositories listedEmploying Large Language Models (LLMs) to address mathematical problems is an intriguing research endeavor, considering the abundance of math problems expressed in natural language across numerous science and…
-
23 May 2023 2 repositories listedRecent developments in large pre-trained language models have enabled unprecedented performance on a variety of downstream tasks.
-
17 May 2022 2 repositories listedTo address this issue and make a step towards interpretable MWP solving, we first construct a high-quality MWP dataset named InterMWP which consists of 11, 495 MWPs and annotates interpretable logical formulas based on…
-
7 Jan 2025 1 repository listedOur approach proposes an adaptive diversity distillation method, in which a student model learns diverse equations by selectively transferring high-quality knowledge from a teacher model.
-
17 Oct 2024 1 repository listedMany students struggle with math word problems (MWPs), often finding it difficult to identify key information and select the appropriate mathematical operations.
-
10 Oct 2024 1 repository listed Syntology ran 8 of 8 samples · 0 unverified · 8 pointer-only (licence)Experiments are conducted on nine benchmarks which demonstrates that our approach improves the reasoning accuracy of LLMs.
-
2 Oct 2024 1 repository listed Syntology ran 0 of 15 samples · 15 unverifiedHowever, most of the cutting-edge progress in mathematical reasoning with LLMs has become \emph{closed-source} due to lack of access to training data.
-
26 Jun 2024 1 repository listed Syntology ran 7 of 12 samples · 5 unverified · 12 pointer-only (licence)Mathematical reasoning presents a significant challenge for Large Language Models (LLMs) due to the extensive and precise chain of reasoning required for accuracy.
-
18 Jun 2024 1 repository listed Syntology ran 10 of 12 samples · 2 unverifiedSolving mathematical problems requires advanced reasoning abilities and presents notable challenges for large language models.
-
6 May 2024 1 repository listedAlthough recent advancements in large language models (LLMs) have significantly improved their performance on various tasks, they still face challenges with complex and symbolic multi-step reasoning, particularly in…
-
23 Apr 2024 1 repository listedTo this end, we propose a simple-yet-effective method, namely Deeply Understanding the Problems (DUP), to improve the LLMs' math problem-solving ability by addressing semantic misunderstanding errors.
-
18 Apr 2024 1 repository listed Syntology ran 13 of 18 samples · 5 unverified · 18 pointer-only (licence)Despite the impressive capabilities of Large Language Models (LLMs) on various tasks, they still struggle with scenarios that involves complex reasoning and planning.
-
6 Apr 2024 1 repository listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Firstly, their effectiveness in tackling complex mathematical problems is somewhat constrained.
-
12 Mar 2024 1 repository listedWe investigate efficient methods for training Large Language Models (LLMs) to possess capabilities in multiple specialized domains, such as coding, math reasoning and world knowledge.
Syntology lines on 18 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections