Papers › MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
Chengpeng Li, Zheng Yuan, Hongyi Yuan, Guanting Dong, Keming Lu, Jiancan Wu, Chuanqi Tan, Xiang Wang, Chang Zhou
In math reasoning with large language models (LLMs), fine-tuning data augmentation by query evolution and diverse reasoning paths is empirically verified effective, profoundly narrowing the gap between open-sourced LLMs and cutting-edge proprietary LLMs. In this paper, we conduct an investigation for such data augmentation in math reasoning and are intended to answer: (1) What strategies of data augmentation are more effective; (2) What is the scaling relationship between the amount of augmented data and model performance; and (3) Can data augmentation incentivize generalization to out-of-domain mathematical reasoning tasks? To this end, we create two new dataset AugGSM8K and AugMATH, by complicating and diversifying the queries and sampling multiple reasoning paths from GSM8K and MATH. We obtained a series of LLMs called MuggleMath by fine-tuning LLaMA models on AugGSM8K and AugMATH. MuggleMath substantially achieves new state-of-the-art on GSM8K and MATH. A log-linear relationship and a segmented log-linear are presented between MuggleMath's performance and the amount of augmented data on GSM8K and MATH, respectively. We also find that it is weak in out-of-domain math reasoning generalization from AugGSM8K to MATH and from AugMATH to GSM8K, which suggests that augmenting queries that cover a broader range of subjects is more beneficial for generalization. We release our codes and augmented data in https://github.com/OFA-Sys/gsm8k-ScRel.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Arithmetic Reasoning | GSM8K | MuggleMATH 70B | Accuracy | 82.3 | #61 of 164 | Archive leaderboard | report |
| Arithmetic Reasoning | GSM8K | MuggleMATH 70B | Parameters (Billion) | 70 | #61 of 164 | Archive leaderboard | report |
| Arithmetic Reasoning | GSM8K | MuggleMATH 13B | Accuracy | 74 | #91 of 164 | Archive leaderboard | report |
| Arithmetic Reasoning | GSM8K | MuggleMATH 13B | Parameters (Billion) | 13 | #91 of 164 | Archive leaderboard | report |
| Arithmetic Reasoning | GSM8K | MuggleMATH 7B | Accuracy | 69.8 | #103 of 164 | Archive leaderboard | report |
| Arithmetic Reasoning | GSM8K | MuggleMATH 7B | Parameters (Billion) | 7 | #103 of 164 | Archive leaderboard | report |
| Math Word Problem Solving | MATH | MuggleMATH-70B | Accuracy | 35.6 | #80 of 135 | Archive leaderboard | report |
| Math Word Problem Solving | MATH | MuggleMATH-70B | Parameters (Billions) | 70 | #80 of 135 | Archive leaderboard | report |
| Math Word Problem Solving | MATH | MuggleMATH-13B | Accuracy | 30.7 | #87 of 135 | Archive leaderboard | report |
| Math Word Problem Solving | MATH | MuggleMATH-13B | Parameters (Billions) | 13 | #87 of 135 | Archive leaderboard | report |
| Math Word Problem Solving | MATH | MuggleMATH 7B | Accuracy | 25.8 | #96 of 135 | Archive leaderboard | report |
| Math Word Problem Solving | MATH | MuggleMATH 7B | Parameters (Billions) | 7 | #96 of 135 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Methods
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections