Browse State-of-the-Art › GSM8K › Papers, page 5
GSM8K
Papers archive 2025-07-28
archive papers tagged: 439 · with a code link: 209 · where Syntology ran a sample: 116 (96 with a run with no instrument failure, 20 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (116 of 439 tagged: 96 with a run with no instrument failure, 20 where every run was a failure of Syntology's instrument)
Page 5 of 5: papers 401 to 439 of 439, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Multi-step Problem Solving Through a Verifier: An Empirical Analysis on Model-induced Process Supervision5 Feb 2024 0 repositories listed
-
YODA: Teacher-Student Progressive Learning for Language Models28 Jan 2024 0 repositories listed
-
Self-Imagine: Effective Unimodal Reasoning with Multimodal Models using Self-Imagination16 Jan 2024 0 repositories listed
-
Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities22 Dec 2023 0 repositories listed
-
From Good to Great: Improving Math Reasoning with Tool-Augmented Interleaf Prompting18 Dec 2023 0 repositories listed
-
14 Dec 2023 0 repositories listed
-
14 Dec 2023 0 repositories listed
-
Training Chain-of-Thought via Latent-Variable Inference28 Nov 2023 0 repositories listed
-
First-Step Advantage: Importance of Starting Right in Multi-Step Math Reasoning14 Nov 2023 0 repositories listed
-
SAIE Framework: Support Alone Isn't Enough -- Advancing LLM Training with Adversarial Remarks14 Nov 2023 0 repositories listed
-
14 Nov 2023 0 repositories listed
-
Let's Reinforce Step by Step10 Nov 2023 0 repositories listed
-
Prompt Engineering a Prompt Engineer9 Nov 2023 0 repositories listed
-
The Alignment Ceiling: Objective Mismatch in Reinforcement Learning from Human Feedback31 Oct 2023 0 repositories listed
-
SEGO: Sequential Subgoal Optimization for Mathematical Problem-Solving19 Oct 2023 0 repositories listed
-
Let's reward step by step: Step-Level reward model as the Navigators for Reasoning16 Oct 2023 0 repositories listed
-
DavIR: Data Selection via Implicit Reward for Large Language Models16 Oct 2023 0 repositories listed
-
11 Oct 2023 0 repositories listed
-
From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference4 Oct 2023 0 repositories listed
-
Large Language Models as Analogical Reasoners3 Oct 2023 0 repositories listed
-
Think before you speak: Training Language Models With Pause Tokens3 Oct 2023 0 repositories listed
-
Adapting LLM Agents with Universal Feedback in Communication1 Oct 2023 0 repositories listed
-
UPAR: A Kantian-Inspired Prompting Framework for Enhancing Large Language Model Capabilities30 Sep 2023 0 repositories listed
-
Contrastive Decoding Improves Reasoning in Large Language Models17 Sep 2023 0 repositories listed
-
Exploring an LM to generate Prolog Predicates from Mathematics Questions7 Sep 2023 0 repositories listed
-
MathAttack: Attacking Large Language Models Towards Math Solving Ability4 Sep 2023 0 repositories listed
-
No Train Still Gain. Unleash Mathematical Reasoning of Large Language Models with Monte Carlo Tree Search Guided by Energy Function1 Sep 2023 0 repositories listed
-
DiversiGATE: A Comprehensive Framework for Reliable Large Language Models22 Jun 2023 0 repositories listed
-
Interpretable Math Word Problem Solution Generation Via Step-by-step Planning1 Jun 2023 0 repositories listed
-
RCOT: Detecting and Rectifying Factual Inconsistency in Reasoning by Reversing Chain-of-Thought19 May 2023 0 repositories listed
-
Hint of Thought prompting: an explainable and zero-shot approach to reasoning tasks with LLMs19 May 2023 0 repositories listed
-
Self-Evaluation Guided Beam Search for Reasoning1 May 2023 0 repositories listed
-
Teaching Small Language Models to Reason16 Dec 2022 0 repositories listed
-
Explicit Knowledge Transfer for Weakly-Supervised Code Generation30 Nov 2022 0 repositories listed
-
25 Nov 2022 0 repositories listed
-
20 Oct 2022 0 repositories listed
-
20 Oct 2022 0 repositories listed
-
Complexity-Based Prompting for Multi-Step Reasoning3 Oct 2022 0 repositories listed
-
6 Jun 2022 0 repositories listed