Datasets › GSM8K › Papers where code ran, page 1
GSM8K
Papers archive 2025-07-28
papers with a benchmark row: 58 · with a code link: 44 · where Syntology ran a sample: 27 (25 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument) Syntology
Show:
all papers with a benchmark row only where code ran (27 of 58 with a benchmark row: 25 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument)
Syntology We ran code from the paper's repository; we did not run it on this dataset or check it against this dataset's benchmarks.
Page 1 of 1: papers 1 to 27 of the 27 papers with a benchmark row here where Syntology ran at least one harvested sample (25 with a run with no instrument failure, 2 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.
A run counts when it is a run with no instrument failure any run, instrument failures included The first hides, in your browser, the papers where every run was a failure of Syntology's instrument; the second shows every paper on this page.
The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 1,881. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since 2026-09-28 , the first build that kept a record date for it; when this build read Syntology's graph is in the build record .
Paper Code Results Date Samples run Syntology
Teaching-Inspired Integrated Prompting Framework: A Novel Approach for Enhancing Reasoning in Large Language Models
1
1
10 Oct 2024
official (archive's flag): 8 ran · 8 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (8 pointer-only for licence )
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data
1
4
2 Oct 2024
12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
1
1
26 Jun 2024
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (12 pointer-only for licence )
Breaking the Ceiling of the LLM Community by Treating Token Generation as a Classification for Ensembling
1
1
18 Jun 2024
official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified (1 pointer-only for licence )
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
1
8
18 Jun 2024
official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing
1
2
18 Apr 2024
official (archive's flag): 14 ran · 14 ran (of which 1 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (18 pointer-only for licence )
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning
1
4
23 Feb 2024
official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (11 pointer-only for licence )
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
5
1
5 Feb 2024
official (archive's flag): 8 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (3 pointer-only for licence )
Frugal LMs Trained to Invoke Symbolic Solvers Achieve Parameter-Efficient Arithmetic Reasoning
1
1
9 Dec 2023
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 1 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (2 pointer-only for licence )
OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning
1
3
16 Nov 2023
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence )
Llemma: An Open Language Model For Mathematics
4
2
16 Oct 2023
official (archive's flag): 6 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
Mistral 7B
6
1
10 Oct 2023
official (archive's flag): 9 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence )
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning
1
6
5 Oct 2023
official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving
1
6
29 Sep 2023
official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (9 pointer-only for licence )
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
1
4
21 Sep 2023
official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 14 where Syntology's instrument failed) · 7 unverified
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
1
4
18 Aug 2023
10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (16 pointer-only for licence )
Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
1
3
3 Aug 2023
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (3 pointer-only for licence )
Llama 2: Open Foundation and Fine-Tuned Chat Models
19
1
18 Jul 2023
community repositories only · 33 ran (of which 7 constructed an object rather than computing a result; 22 with no instrument failure: 1 honoured, 1 violated, 20 with no contract checked; 11 where Syntology's instrument failed) · 19 unverified (20 pointer-only for licence )
CodeT5+: Open Code Large Language Models for Code Understanding and Generation
2
1
13 May 2023
official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified
Sparks of Artificial General Intelligence: Early experiments with GPT-4
3
1
22 Mar 2023
9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
GPT-4 Technical Report
11
1
15 Mar 2023
community repositories only · 5 ran (of which 0 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 2 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence )
LLaMA: Open and Efficient Foundation Language Models
57
8
27 Feb 2023
official: no sample here; runs from other or unrecorded repositories · 37 ran (of which 9 constructed an object rather than computing a result; 25 with no instrument failure: 3 honoured, 0 violated, 22 with no contract checked; 12 where Syntology's instrument failed) · 21 unverified (4 pointer-only for licence )
LEVER: Learning to Verify Language-to-Code Generation with Execution
1
1
16 Feb 2023
official (archive's flag): 18 ran · 18 ran (of which 0 constructed an object rather than computing a result; 18 with no instrument failure: 0 honoured, 1 violated, 17 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (1 pointer-only for licence )
Learning Math Reasoning from Self-Sampled Correct and Partially-Correct Solutions
1
2
28 May 2022
official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified
Large Language Models are Zero-Shot Reasoners
4
7
24 May 2022
official: harvested, nothing ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (1 pointer-only for licence )
UL2: Unifying Language Learning Paradigms
2
2
10 May 2022
community repositories only · 15 ran (of which 0 constructed an object rather than computing a result; 15 with no instrument failure: 0 honoured, 0 violated, 15 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
Self-Consistency Improves Chain of Thought Reasoning in Language Models
3
1
21 Mar 2022
1 ran (of which 1 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 0 violated, 1 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; the one sample that ran constructed an object rather than computing a result