Papers › Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

14 Dec 2023arXiv:2312.08935archive 2025-07-28

Peiyi Wang, Lei LI, Zhihong Shao, R. X. Xu, Damai Dai, Yifei Li, Deli Chen, Y. Wu, Zhifang Sui

In this paper, we present an innovative process-oriented math process reward model called \textbf{Math-Shepherd}, which assigns a reward score to each step of math problem solutions. The training of Math-Shepherd is achieved using automatically constructed process-wise supervision data, breaking the bottleneck of heavy reliance on manual annotation in existing work. We explore the effectiveness of Math-Shepherd in two scenarios: 1) \textit{Verification}: Math-Shepherd is utilized for reranking multiple outputs generated by Large Language Models (LLMs); 2) \textit{Reinforcement Learning}: Math-Shepherd is employed to reinforce LLMs with step-by-step Proximal Policy Optimization (PPO). With Math-Shepherd, a series of open-source LLMs demonstrates exceptional performance. For instance, the step-by-step PPO with Math-Shepherd significantly improves the accuracy of Mistral-7B (77.9\%→84.1\% on GSM8K and 28.6\%→33.0\% on MATH). The accuracy can be further enhanced to 89.1\% and 43.5\% on GSM8K and MATH with the verification of Math-Shepherd, respectively. We believe that automatic process supervision holds significant potential for the future evolution of LLMs.

PaperPDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

hkust-nlp/b-star mentioned on GitHubpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Arithmetic ReasoningGSM8KMathMath Word Problem SolvingMathematical ReasoningReranking

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Arithmetic Reasoning GSM8K Shepherd+Mistral-7B (SFT on MetaMATH + PRM RL+ PRM rerank, k=256) Accuracy 89.1 #24 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Shepherd+Mistral-7B (SFT on MetaMATH + PRM RL+ PRM rerank, k=256) Parameters (Billion) 7 #24 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Shepherd + Mistral-7B (SFT on MetaMATH + PRM RL) Accuracy 84.1 #52 of 164 Archive leaderboard report
Arithmetic Reasoning GSM8K Shepherd + Mistral-7B (SFT on MetaMATH + PRM RL) Parameters (Billion) 7 #52 of 164 Archive leaderboard report
Math Word Problem Solving MATH Shepherd + DeepSeek-67B (SFT on MetaMATH + PRM rerank, k=256) Accuracy 48.1 #55 of 135 Archive leaderboard report
Math Word Problem Solving MATH Shepherd + DeepSeek-67B (SFT on MetaMATH + PRM rerank, k=256) Parameters (Billions) 67 #55 of 135 Archive leaderboard report
Math Word Problem Solving MATH Shepherd+Mistral-7B (SFT on MetaMATH + PRM RL+ PRM rerank, k=256) Accuracy 43.5 #71 of 135 Archive leaderboard report
Math Word Problem Solving MATH Shepherd+Mistral-7B (SFT on MetaMATH + PRM RL+ PRM rerank, k=256) Parameters (Billions) 7 #71 of 135 Archive leaderboard report
Math Word Problem Solving MATH Shepherd + Mistral-7B (SFT on MetaMATH + PRM RL) Accuracy 33.0 #84 of 135 Archive leaderboard report
Math Word Problem Solving MATH Shepherd + Mistral-7B (SFT on MetaMATH + PRM RL) Parameters (Billions) 7 #84 of 135 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Entropy RegularizationPPO

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections