Datasets › MATH › Papers where code ran, page 1

MATH

Papers archive 2025-07-28

papers with a benchmark row: 37 · with a code link: 34 · where Syntology ran a sample: 21 (17 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument) Syntology

Show: all papers with a benchmark rowonly where code ran (21 of 37 with a benchmark row: 17 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument)

Syntology We ran code from the paper's repository; we did not run it on this dataset or check it against this dataset's benchmarks.

Page 1 of 1: papers 1 to 21 of the 21 papers with a benchmark row here where Syntology ran at least one harvested sample (17 with a run with no instrument failure, 4 where every run was a failure of Syntology's instrument), newest first by the archive's date (ties by arXiv id). This is a filter on Syntology's record ordered by date only, not a ranking; a run is not a correctness claim. A paper missing from this list is not a recorded non-run: it may have no arXiv id, no harvested code, or only samples that have not run yet.

The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset, not that list; the archive's count for this dataset is 1,330. The Syntology column is from Syntology's graph, stated per sample; it is not part of any archive number. A line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified” (C is a part of N, never taken away from it); the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. When the archive marks a repository official for the paper, the cell starts with that repository's state (the archive's flag, not a verdict on who wrote the code; “community repositories only” when every sample that ran came from a community repository, “official: no sample here; runs from other or unrecorded repositories” when some came from a repository the paper names or has in its text, or from none recorded); hover it for the repositories the samples that ran came from. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record.

PaperCodeResultsDateSamples run Syntology
OpenMathInstruct-2: Accelerating AI for Math with Massive Open-Source Instruction Data 1 4 2 Oct 2024 12 ran (of which 0 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 0 violated, 12 with no contract checked; 0 where Syntology's instrument failed) · 3 unverified
Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs 1 1 26 Jun 2024 official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (12 pointer-only for licence)
DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving 1 8 18 Jun 2024 official (archive's flag): 10 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 2 unverified
Toward Self-Improvement of LLMs via Imagination, Searching, and Criticizing 1 1 18 Apr 2024 official (archive's flag): 14 ran · 14 ran (of which 1 constructed an object rather than computing a result; 12 with no instrument failure: 0 honoured, 1 violated, 11 with no contract checked; 2 where Syntology's instrument failed) · 4 unverified (18 pointer-only for licence)
An Empirical Study of Data Ability Boundary in LLMs' Math Reasoning 1 4 23 Feb 2024 official (archive's flag): 9 ran · 9 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 2 where Syntology's instrument failed) · 2 unverified (11 pointer-only for licence)
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models 5 2 5 Feb 2024 official (archive's flag): 8 ran · 14 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 4 where Syntology's instrument failed) · 10 unverified (3 pointer-only for licence)
Augmenting Math Word Problems via Iterative Question Composing 1 1 17 Jan 2024 official (archive's flag): 4 ran · 4 ran (of which 0 constructed an object rather than computing a result; 4 with no instrument failure: 0 honoured, 0 violated, 4 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified (4 pointer-only for licence)
Mixtral of Experts 6 2 8 Jan 2024 5 ran (of which 5 constructed an object rather than computing a result; 5 with no instrument failure: 0 honoured, 0 violated, 5 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified; every one of the 5 samples that ran constructed an object rather than computing a result
Mistral 7B 6 1 10 Oct 2023 official (archive's flag): 9 ran · 10 ran (of which 0 constructed an object rather than computing a result; 10 with no instrument failure: 0 honoured, 0 violated, 10 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (2 pointer-only for licence)
MathCoder: Seamless Code Integration in LLMs for Enhanced Mathematical Reasoning 1 6 5 Oct 2023 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified
ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving 1 8 29 Sep 2023 official (archive's flag): 11 ran · 11 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (9 pointer-only for licence)
MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models 1 3 21 Sep 2023 official (archive's flag): 15 ran · 15 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 14 where Syntology's instrument failed) · 7 unverified
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct 1 4 18 Aug 2023 10 ran (of which 0 constructed an object rather than computing a result; 8 with no instrument failure: 0 honoured, 0 violated, 8 with no contract checked; 2 where Syntology's instrument failed) · 6 unverified (16 pointer-only for licence)
Cumulative Reasoning with Large Language Models 1 2 8 Aug 2023 official (archive's flag): 5 ran · 5 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 2 where Syntology's instrument failed) · 1 unverified (6 pointer-only for licence)
Let's Verify Step by Step 3 1 31 May 2023 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 1 honoured, 2 violated, 0 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
Progressive-Hint Prompting Improves Reasoning in Large Language Models 1 1 19 Apr 2023 official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (4 pointer-only for licence)
Sparks of Artificial General Intelligence: Early experiments with GPT-4 3 1 22 Mar 2023 9 ran (of which 0 constructed an object rather than computing a result; 9 with no instrument failure: 0 honoured, 0 violated, 9 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified
LLaMA: Open and Efficient Foundation Language Models 57 8 27 Feb 2023 official: no sample here; runs from other or unrecorded repositories · 37 ran (of which 9 constructed an object rather than computing a result; 25 with no instrument failure: 3 honoured, 0 violated, 22 with no contract checked; 12 where Syntology's instrument failed) · 21 unverified (4 pointer-only for licence)
PAL: Program-aided Language Models 3 1 18 Nov 2022 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified
Galactica: A Large Language Model for Science 1 7 16 Nov 2022 official (archive's flag): 2 ran · 2 ran (of which 0 constructed an object rather than computing a result; 2 with no instrument failure: 0 honoured, 0 violated, 2 with no contract checked; 0 where Syntology's instrument failed) · 0 unverified
Measuring Mathematical Problem Solving With the MATH Dataset 5 8 5 Mar 2021 official (archive's flag): 2 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 0 unverified (1 pointer-only for licence)