Browse › Reasoning › Arithmetic Reasoning › GSM8K

Arithmetic Reasoning archive 2025-07-28

GSM8K Benchmark (Arithmetic Reasoning)

164 rows 118 with code listed 2 metrics Dataset page

Over time archive 2025-07-28

The chart needs JavaScript; the table below carries every value.

Direction inferred from the metric name, not from the archive: Accuracy (higher is better). Not inferred (points only, no best-so-far line): Parameters (Billion). Points are placed at the row's paper date; 150 of 164 rows carry one.

Results archive 2025-07-28

Archive rows end at the archive snapshot, 2025-07-28: no result published after that date is in this table. Rank is the archive's row order at that snapshot; not re-ranked here. Metric values are the archive's strings. Column headers sort the table in your browser; each row keeps its archive rank.

Paper Code Ran Syntology Report
1 Claude 3.5 Sonnet (HPT) 97.72 – Paper Code 2024 linked, not harvested report
2 DUP prompt upon GPT-4 97.1 – Paper Code 2024 linked, not harvested report
3 Qwen2-Math-72B-Instruct (greedy) 96.772 ✓ Paper Code 2024 linked, not harvested report
4 SFT-Mistral-7B (Metamath, OVM, Smart Ensemble) 96.47 ✓ – – not matched report
5 OpenMath2-Llama3.1-70B (majority@256) 96.0 ✓ Paper Code 2024 12 of 15 ran · 3 unverified report
6 Jiutian-大模型 95.275 – – – not matched report
7 DAMOMath-7B(MetaMath, OVM, BS, Ensemble) 95.17 ✓ – – not matched report
8 Claude 3 Opus (0-shot chain-of-thought) 95 – Paper – 2024 no code linked report
9 OpenMath2-Llama3.1-70B 94.9 ✓ Paper Code 2024 12 of 15 ran · 3 unverified report
10 GPT-4 (Teaching-Inspired) 94.8 – Paper Code 2024 8 of 8 ran · 0 unverified report
11 SFT-Mistral-7B (Metamath + ovm +ensemble) 94.137 ✓ – – not matched report
12 OpenMath2-Llama3.1-8B (majority@256) 94.1 ✓ Paper Code 2024 12 of 15 ran · 3 unverified report
13 Qwen2-72B-Instruct-Step-DPO (0-shot CoT) 94.0 ✓ Paper Code 2024 7 of 12 ran · 5 unverified report
14 DAMOMath-7B(MetaMath, OVM, Ensemble) 93.27 ✓ – – not matched report
15 Claude 3 Sonnet (0-shot chain-of-thought) 92.3 – Paper – 2024 no code linked report
16 AlphaLLM (with MCTS) 9270 – Paper Code 2024 13 of 18 ran · 5 unverified report
17 OpenMath2-Llama3.1-8B 91.7 ✓ Paper Code 2024 12 of 15 ran · 3 unverified report
18 PaLM 2 (few-shot, k=8, SC) 91.0 – Paper Code 2023 linked, not harvested report
19 GaC(Qwen2-72B-Instruct + Llama-3-70B-Instruct) 90.91 – Paper Code 2024 4 of 7 ran · 3 unverified report
20 OpenMath-CodeLlama-70B (w/ code, SC, k=50) 90.870 ✓ Paper Code 2024 linked, not harvested report
21 DART-Math-Llama3-70B-Uniform (0-shot CoT, w/o code) 90.470 ✓ Paper Code 2024 10 of 12 ran · 2 unverified report
22 OpenMath-Llama2-70B (w/ code, SC, k=50) 90.170 ✓ Paper Code 2024 linked, not harvested report
23 DART-Math-Llama3-70B-Prop2Diff (0-shot CoT, w/o code) 89.670 ✓ Paper Code 2024 10 of 12 ran · 2 unverified report
24 Shepherd+Mistral-7B (SFT on MetaMATH + PRM RL+ PRM rerank, k=256) 89.17 ✓ Paper Code 2023 linked, not harvested report
25 Llama SFT (Metamath ToRA Ensemble) 89.013 ✓ – – not matched report
26 Minerva 62B (maj5@100) 8962 – Paper Code 2022 linked, not harvested report
27 Claude 3 Haiku (0-shot chain-of-thought) 88.9 – Paper – 2024 no code linked report
28 ToRA-70B (SC, k=50) 88.370 ✓ Paper Code 2023 6 of 12 ran · 6 unverified report
29 DeepSeekMATH-RL-7B 88.27 ✓ Paper Code 2024 10 of 24 ran · 14 unverified report
30 DART-Math-DSMath-7B-Uniform (0-shot CoT, w/o code) 88.27 ✓ Paper Code 2024 10 of 12 ran · 2 unverified report
31 OpenMath-CodeLlama-34B (w/ code, SC, k=50) 88.034 ✓ Paper Code 2024 linked, not harvested report
32 Claude 2 (0-shot chain-of-thought) 88 – Paper – 2023 no code linked report
33 Shivaay-4B (8-shot chain-of-thought) 87.414 – – – not matched report
34 DeepMind 70B Model (SFT+ORM-RL, ORM reranking) 87.370 ✓ Paper – 2022 no code linked report
35 MMOS-DeepSeekMath-7B(0-shot,k=50) 87.27 ✓ Paper Code 2024 9 of 11 ran · 2 unverified report
36 DeepMind 70B Model (SFT+PRM-RL, PRM reranking) 87.170 ✓ Paper – 2022 no code linked report
37 GPT-4 87.1 – Paper Code 2023 9 of 10 ran · 1 unverified report
38 OpenMath-Mistral-7B (w/ code, SC, k=50) 86.97 ✓ Paper Code 2024 linked, not harvested report
39 Orca-Math 7B (fine-tuned) 86.87 ✓ Paper – 2024 no code linked report
40 DART-Math-DSMath-7B-Prop2Diff (0-shot CoT, w/o code) 86.87 ✓ Paper Code 2024 10 of 12 ran · 2 unverified report
41 OpenMath-CodeLlama-13B (w/ code, SC, k=50) 86.813 ✓ Paper Code 2024 linked, not harvested report
42 Gemini Pro (maj1@32) 86.5 – Paper Code 2023 linked, not harvested report
43 Codex (Self-Evaluation Guided Decoding, PAL, multiple reasoning chains, 9-shot gen, 5-shot eval) 85.5 – – – not matched report
44 Claude 1.3 (0-shot chain-of-thought) 85.2 – Paper – 2023 no code linked report
45 ToRA-Code-34B (SC, k=50) 85.134 ✓ Paper Code 2023 6 of 12 ran · 6 unverified report
46 OpenMath-CodeLlama-7B (w/ code, SC, k=50) 84.87 ✓ Paper Code 2024 linked, not harvested report
47 OVM-Mistral-7B (verify100@1) 84.77 – Paper Code 2023 3 of 3 ran · 0 unverified report
48 OpenMath-Llama2-70B (w/ code) 84.770 ✓ Paper Code 2024 linked, not harvested report
49 OpenMath-CodeLlama-70B (w/ code) 84.670 ✓ Paper Code 2024 linked, not harvested report
50 code-davinci-002 175B (LEVER, 8-shot) 84.5175 – Paper Code 2023 18 of 22 ran · 4 unverified report
51 ToRA 70B 84.370 ✓ Paper Code 2023 6 of 12 ran · 6 unverified report
52 Shepherd + Mistral-7B (SFT on MetaMATH + PRM RL) 84.17 ✓ Paper Code 2023 linked, not harvested report
53 MathCoder-L-70B 83.970 ✓ Paper Code 2023 2 of 2 ran · 0 unverified report
54 WizardMath-7B-V1.1 83.27 ✓ Paper Code 2023 10 of 16 ran · 6 unverified report
55 DIVERSE 175B (8-shot) 83.2175 – Paper – 2022 no code linked report
56 OVM-Mistral-7B (verify20@1) 82.67 – Paper Code 2023 3 of 3 ran · 0 unverified report
57 DART-Math-Mistral-7B-Uniform (0-shot CoT, w/o code) 82.67 ✓ Paper Code 2024 10 of 12 ran · 2 unverified report
58 ChatGPT (Ask, Refine, Trust) 82.6 – Paper – 2023 no code linked report
59 DART-Math-Llama3-8B-Uniform (0-shot CoT, w/o code) 82.58 ✓ Paper Code 2024 10 of 12 ran · 2 unverified report
60 MetaMath 70B 82.370 ✓ Paper Code 2023 15 of 22 ran · 7 unverified report
61 MuggleMATH 70B 82.370 ✓ Paper Code 2023 linked, not harvested report
62 PaLM 540B (Self Improvement, Self Consistency) 82.1540 – Paper – 2022 no code linked report
63 MathCoder-CL-34B 81.734 ✓ Paper Code 2023 2 of 2 ran · 0 unverified report
64 WizardMath-70B-V1.0 81.670 ✓ Paper Code 2023 10 of 16 ran · 6 unverified report
65 Phi-GSM+V 1.3B+1.3B (verify48@1) 81.52.6 – Paper – 2023 no code linked report
66 DART-Math-Mistral-7B-Prop2Diff (0-shot CoT, w/o code) 81.17 ✓ Paper Code 2024 10 of 12 ran · 2 unverified report
67 DART-Math-Llama3-8B-Prop2Diff (0-shot CoT, w/o code) 81.18 ✓ Paper Code 2024 10 of 12 ran · 2 unverified report
68 Claude Instant 1.1 (0-shot chain-of-thought) 80.9 – Paper – 2023 no code linked report
69 ToRA-Code 34B 80.734 ✓ Paper Code 2023 6 of 12 ran · 6 unverified report
70 OpenMath-CodeLlama-34B (w/ code) 80.734 ✓ Paper Code 2024 linked, not harvested report
71 PaLM 2 (few-shot, k=8, CoT) 80.7 – Paper Code 2023 linked, not harvested report
72 MMOS-DeepSeekMath-7B(0-shot) 80.57 ✓ Paper Code 2024 9 of 11 ran · 2 unverified report
73 MMOS-CODE-34B(0-shot) 80.434 ✓ Paper Code 2024 9 of 11 ran · 2 unverified report
74 OpenMath-Mistral-7B (w/ code) 80.27 ✓ Paper Code 2024 linked, not harvested report
75 Self-Evaluation Guided Decoding (Codex, PAL, single reasoning chain, 9-shot gen, 5-shot eval) 80.2 – – – not matched report
76 OpenMath-CodeLlama-13B (w/ code) 78.813 ✓ Paper Code 2024 linked, not harvested report
77 Minerva 540B (CoT) 78.5540 – Paper Code 2022 linked, not harvested report
78 Camelidae-8×34B (5-shot) 78.3 – Paper Code 2024 linked, not harvested report
79 Qwen2idae-16x14B (5-shot) 77.8 – Paper Code 2024 linked, not harvested report
80 MetaMath-Mistral-7B 77.77 ✓ Paper Code 2023 15 of 22 ran · 7 unverified report
81 OpenChat-3.5 7B 77.37 – Paper Code 2023 0 of 1 ran · 1 unverified report
82 DeepMind 70B Model (STaR, maj1@96) 76.570 ✓ Paper – 2022 no code linked report
83 Arithmo2-Mistral-7B 76.47 – – – not matched report
84 OpenMath-CodeLlama-7B (w/ code) 75.97 ✓ Paper Code 2024 linked, not harvested report
85 ToRA-Code 13B 75.813 ✓ Paper Code 2023 6 of 12 ran · 6 unverified report
86 Arithmo-Mistral-7B 74.77 – – – not matched report
87 PaLM 540B maj1@40 (8-shot) 74.4540 ✓ Paper Code 2022 1 of 1 ran · 0 unverified report
88 PaLM 540B (Self Consistency) 74.4540 – Paper – 2022 no code linked report
89 Phi-GSM 2.7B (fine-tuned) 74.32.7 – Paper – 2023 no code linked report
90 MathCoder-CL-13B 74.17 ✓ Paper Code 2023 2 of 2 ran · 0 unverified report
91 MuggleMATH 13B 7413 ✓ Paper Code 2023 linked, not harvested report
92 MMOS-CODE-7B(0-shot) 73.97 ✓ Paper Code 2024 9 of 11 ran · 2 unverified report
93 CodeT5+ 73.80.77 – Paper Code 2023 3 of 4 ran · 1 unverified report
94 Llama-3.3-70B + CAPO 73.73 – Paper Code 2025 linked, not harvested report
95 OVM-Llama2-7B (verify100@1) 73.77 – Paper Code 2023 3 of 3 ran · 0 unverified report
96 PaLM 540B (Self Improvement, CoT Prompting) 73.5540 – Paper – 2022 no code linked report
97 KwaiYiiMath 13B 73.313 ✓ Paper – 2023 no code linked report
98 ToRA-Code 7B 72.67 ✓ Paper Code 2023 6 of 12 ran · 6 unverified report
99 MathCoder-L-13B 72.613 ✓ Paper Code 2023 2 of 2 ran · 0 unverified report
100 DBRX Base 132B 72.3 – – – not matched report
101 Self-Evaluation Guided Decoding (Codex, CoT, single reasoning chain, 9-shot gen, 5-shot eval) 71.9 – – – not matched report
102 MetaMath 13B 71.013 ✓ Paper Code 2023 15 of 22 ran · 7 unverified report
103 MuggleMATH 7B 69.87 ✓ Paper Code 2023 linked, not harvested report
104 LLaMA 65B-maj1@k 69.765 – Paper Code 2023 37 of 58 ran · 21 unverified report
105 Minerva 62B (maj1@100) 68.562 ✓ Paper Code 2022 linked, not harvested report
106 code-davinci-002 (Least-to-Most Prompting) 68.01175 – Paper Code 2022 linked, not harvested report
107 MathCoder-CL-7B 67.87 ✓ Paper Code 2023 2 of 2 ran · 0 unverified report
108 DBRX Instruct 132B 66.9 – – – not matched report
109 MetaMath 7B 66.47 ✓ Paper Code 2023 15 of 22 ran · 7 unverified report
110 Mistral-Small-24B + CAPO 65.07 – Paper Code 2025 linked, not harvested report
111 RFT 70B 64.879 ✓ Paper Code 2023 3 of 3 ran · 0 unverified report
112 MathCoder-L-7B 64.27 ✓ Paper Code 2023 2 of 2 ran · 0 unverified report
113 WizardMath-13B-V1.0 63.913 ✓ Paper Code 2023 10 of 16 ran · 6 unverified report
114 GPT-J (CoRe) 63.212 – Paper Code 2022 linked, not harvested report
115 Llama-2 70B (on 100 first questions, 4-shot, auto-optimized prompting) 6170 – Paper – 2024 no code linked report
116 Qwen2.5-32B + CAPO 60.2 – Paper Code 2025 linked, not harvested report
117 LLaMA 2 70B (CoT-Influx) 59.5970 – Paper – 2023 no code linked report
118 Orca 2 13B 59.1413 – Paper – 2023 no code linked report
119 U-PaLM 58.5540 – Paper – 2022 no code linked report
120 PaLM-540B (few-Shot-cot) 58.1540 ✓ Paper Code 2022 3 of 4 ran · 1 unverified report
121 GPT-3.5 (few-shot, k=5) 57.1 – Paper Code 2023 5 of 5 ran · 0 unverified report
122 Minerva 8B (maj5@100) 56.88 – Paper Code 2022 linked, not harvested report
123 LLaMA 2 70B (on-shot) 56.870 – Paper Code 2023 31 of 52 ran · 21 unverified report
124 PaLM 540B (8-shot) 56.5540 ✓ Paper Code 2022 linked, not harvested report
125 PaLM 540B (CoT Prompting) 56.5540 – Paper – 2022 no code linked report
126 RFT 13B 55.313 ✓ Paper Code 2023 3 of 3 ran · 0 unverified report
127 Finetuned GPT-3 175B + verifier 55.0175 ✓ Paper Code 2022 3 of 4 ran · 1 unverified report
128 WizardMath-7B-V1.0 54.97 ✓ Paper Code 2023 10 of 16 ran · 6 unverified report
129 LLaMA 33B-maj1@k 53.133 – Paper Code 2023 37 of 58 ran · 21 unverified report
130 Minerva 62B (8-shot) 52.462 ✓ Paper Code 2022 linked, not harvested report
131 Mistral 7B (maj@8) 52.27 – Paper Code 2023 10 of 11 ran · 1 unverified report
132 Llemma 34B 51.534 – Paper Code 2023 6 of 8 ran · 2 unverified report
133 Text-davinci-002-175B (zero-plus-few-Shot-cot (8 samples)) 51.5175 ✓ Paper Code 2022 3 of 4 ran · 1 unverified report
134 RFT 7B 51.27 ✓ Paper Code 2023 3 of 3 ran · 0 unverified report
135 LLaMA 65B 50.965 – Paper Code 2023 37 of 58 ran · 21 unverified report
136 Orca 2 7B 47.237 – Paper – 2023 no code linked report
137 Llama-2 13B (on 100 first questions, 4-shot, auto-optimized prompting) 4313 – Paper – 2024 no code linked report
138 text-davinci-002 175B (2-shot, CoT) 41.3175 ✓ Paper Code 2022 3 of 4 ran · 1 unverified report
139 Mistral 7B (on 100 first questions, 4-shot, auto-optimized prompting) 417 – Paper – 2024 no code linked report
140 text-davinci-002 175B (0-shot, CoT) 40.7175 ✓ Paper Code 2022 3 of 4 ran · 1 unverified report
141 Branch-Train-MiX 4x7B (sampling top-2 experts) 37.1 – Paper Code 2024 linked, not harvested report
142 Llemma 7B 36.47 – Paper Code 2023 6 of 8 ran · 2 unverified report
143 LLaMA 33B 35.633 – Paper Code 2023 37 of 58 ran · 21 unverified report
144 Vicuna (SYRELM) 35.213 ✓ Paper Code 2023 2 of 2 ran · 0 unverified report
145 PaLM 62B (8-shot) 33.062 ✓ Paper Code 2022 linked, not harvested report
146 PaLM 540B (Self Improvement, Standard-Prompting) 32.2540 – Paper – 2022 no code linked report
147 LLaMA 13B-maj1@k 29.313 – Paper Code 2023 37 of 58 ran · 21 unverified report
148 Minerva 8B-maj1@k (8-shot) 28.48 ✓ Paper Code 2022 linked, not harvested report
149 GPT-2-Medium 355M + question-solution classifier (BS=5) 20.80.355 – Paper – 2022 no code linked report
150 GPT-Neo-2.7B + Self-Sampling 19.52.7 – Paper Code 2022 7 of 11 ran · 4 unverified report
151 GPT-2-Medium 355M (fine-tuned, BS=5) 18.30.355 – Paper – 2022 no code linked report
152 LLaMA 7B (maj1@k) 18.17 – Paper Code 2023 37 of 58 ran · 21 unverified report
153 PaLM 540B (few-shot) 17.9540 ✓ Paper Code 2022 3 of 4 ran · 1 unverified report
154 PaLM 540B (Standard-Prompting) 17.9540 – Paper – 2022 no code linked report
155 LLaMA 13B 17.813 – Paper Code 2023 37 of 58 ran · 21 unverified report
156 GPT-2-Medium 355M + question-solution classifier (BS=1) 16.80.355 – Paper – 2022 no code linked report
157 Minerva 8B (8-shot) 16.28 ✓ Paper Code 2022 linked, not harvested report
158 GPT-2-Medium 355M (BS=5) 12.20.355 – Paper – 2022 no code linked report
159 LLaMA 7B 11.07 – Paper Code 2023 37 of 58 ran · 21 unverified report
160 Text-davinci-002-175B (0-shot) 10.4175 ✓ Paper Code 2022 3 of 4 ran · 1 unverified report
161 GPT-Neo 125M + Self-Sampling 7.50.125 – Paper Code 2022 7 of 11 ran · 4 unverified report
162 UL2 20B (chain-of-thought) 4.420 – Paper Code 2022 15 of 16 ran · 1 unverified report
163 PaLM 8B (8-shot) 4.18 ✓ Paper Code 2022 linked, not harvested report
164 UL2 20B (0-shot) 4.120 – Paper Code 2022 15 of 16 ran · 1 unverified report

All 164 rows shown. 150 link to a paper page on this site; 85 are marked as using additional training data in the archive. No GitHub stars are tracked; "Code" is the first repository the archive lists for the row. The archive carries no row tags, review links or community-submitted rows for this table; none are shown. archive 2025-07-28

Syntology Ran reads "N of M ran · U unverified": of the M code samples Syntology harvested from repositories linked to that row's paper (joined by arXiv id), N executed on a synthesized input and the other U = M−N are unverified (harvested, no recorded run). It counts code from repositories linked to that row's paper, not this result: the row's number was not reproduced and nothing here is a correctness claim. The other cell texts mean no graph line for the row: "linked, not harvested" (the archive links code, Syntology has not harvested it), "no code linked" (no code link in the archive), "not matched" (the row's paper URL matched no paper on this site). 77 rows have a graph line, from 28 distinct papers; 76 rows (27 papers) have at least one sample that ran. Counting each paper once: Syntology ran 259 of 370 samples; 111 unverified. Separately, 96 of those 370 are pointer-only (licence): the site points at that code rather than redistributing it, a licence property recorded for ran and unverified samples alike; each cell's tooltip carries the row's own pointer-only count. Read from the graph 2026-09-25. Per-sample status is on the paper page.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections