Browse State-of-the-Art › GSM8K › Papers, page 4
GSM8K
Papers archive 2025-07-28
archive papers tagged: 439 · with a code link: 209 · where Syntology ran a sample: 116 (96 with a run with no instrument failure, 20 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (116 of 439 tagged: 96 with a run with no instrument failure, 20 where every run was a failure of Syntology's instrument)
Page 4 of 5: papers 301 to 400 of 439, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Evolutionary Pre-Prompt Optimization for Mathematical Reasoning5 Dec 2024 0 repositories listed
-
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning4 Dec 2024 0 repositories listed
-
MALT: Improving Reasoning with Multi-Agent LLM Training2 Dec 2024 0 repositories listed
-
Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference27 Nov 2024 0 repositories listed
-
Predicting Emergent Capabilities by Finetuning25 Nov 2024 0 repositories listed
-
Unraveling Arithmetic in Large Language Models: The Role of Algebraic Structures25 Nov 2024 0 repositories listed
-
Patience Is The Key to Large Language Model Reasoning20 Nov 2024 0 repositories listed
-
Adaptive Decoding via Latent Preference Optimization14 Nov 2024 0 repositories listed
-
Dynamic Subset Tuning: Expanding the Operational Range of Parameter-Efficient Training for Large Language Models13 Nov 2024 0 repositories listed
-
Quasi-random Multi-Sample Inference for Large Language Models9 Nov 2024 0 repositories listed
-
Reasoning Robustness of LLMs to Adversarial Typographical Errors8 Nov 2024 0 repositories listed
-
Kwai-STaR: Transform LLMs into State-Transition Reasoners7 Nov 2024 0 repositories listed
-
Self-Consistency Preference Optimization6 Nov 2024 0 repositories listed
-
Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models2 Nov 2024 0 repositories listed
-
Rethinking Data Synthesis: A Teacher Model Training Recipe with Interpretation27 Oct 2024 0 repositories listed
-
ReasonAgain: Using Extractable Symbolic Programs to Evaluate Mathematical Reasoning24 Oct 2024 0 repositories listed
-
Adaptive Dense Reward: Understanding the Gap Between Action and Reward Space in Alignment23 Oct 2024 0 repositories listed
-
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation22 Oct 2024 0 repositories listed
-
On Designing Effective RL Reward at Training Time for LLM Reasoning19 Oct 2024 0 repositories listed
-
TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling18 Oct 2024 0 repositories listed
-
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs15 Oct 2024 0 repositories listed
-
Nudging: Inference-time Alignment of LLMs via Guided Decoding11 Oct 2024 0 repositories listed
-
Towards Multilingual LLM Evaluation for European Languages11 Oct 2024 0 repositories listed
-
Dialectical Behavior Therapy Approach to LLM Prompting10 Oct 2024 0 repositories listed
-
Think Beyond Size: Adaptive Prompting for More Effective Reasoning10 Oct 2024 0 repositories listed
-
Subtle Errors Matter: Preference Learning via Error-injected Self-editing9 Oct 2024 0 repositories listed
-
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning8 Oct 2024 0 repositories listed
-
PortLLM: Personalizing Evolving Large Language Models with Training-Free and Portable Model Patches8 Oct 2024 0 repositories listed
-
Reasoning Paths Optimization: Learning to Reason and Explore From Diverse Paths7 Oct 2024 0 repositories listed
-
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification5 Oct 2024 0 repositories listed
-
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation3 Oct 2024 0 repositories listed
-
BrainTransformers: SNN-LLM3 Oct 2024 0 repositories listed
-
CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoning3 Oct 2024 0 repositories listed
-
The Role of Deductive and Inductive Reasoning in Large Language Models3 Oct 2024 0 repositories listed
-
Unlocking Structured Thinking in Language Models with Cognitive Prompting3 Oct 2024 0 repositories listed
-
PersonaMath: Enhancing Math Reasoning through Persona-Driven Data Augmentation2 Oct 2024 0 repositories listed
-
Instance-adaptive Zero-shot Chain-of-Thought Prompting30 Sep 2024 0 repositories listed
-
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ25 Sep 2024 0 repositories listed
-
PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning25 Sep 2024 0 repositories listed
-
Uncovering Latent Chain of Thought Vectors in Language Models21 Sep 2024 0 repositories listed
-
ControlMath: Controllable Data Generation Promotes Math Generalist Models20 Sep 2024 0 repositories listed
-
18 Sep 2024 0 repositories listed
-
CPL: Critical Plan Step Learning Boosts LLM Generalization in Reasoning Tasks13 Sep 2024 0 repositories listed
-
STUN: Structured-Then-Unstructured Pruning for Scalable MoE Pruning10 Sep 2024 0 repositories listed
-
Strategic Chain-of-Thought: Guiding Accurate Reasoning in LLMs through Strategy Elicitation5 Sep 2024 0 repositories listed
-
Building Math Agents with Multi-Turn Iterative Preference Learning4 Sep 2024 0 repositories listed
-
Prompt Baking4 Sep 2024 0 repositories listed
-
S³c-Math: Spontaneous Step-level Self-correction Makes Large Language Models Better Mathematical Reasoners3 Sep 2024 0 repositories listed
-
Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic29 Aug 2024 0 repositories listed
-
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems29 Aug 2024 0 repositories listed
-
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models28 Aug 2024 0 repositories listed
-
Threshold Filtering Packing for Supervised Fine-Tuning: Training Related Samples within Packs18 Aug 2024 0 repositories listed
-
SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models16 Aug 2024 0 repositories listed
-
Concise Thoughts: Impact of Output Length on LLM Reasoning and Cost29 Jul 2024 0 repositories listed
-
Cool-Fusion: Fuse Large Language Models without Training29 Jul 2024 0 repositories listed
-
Reliable Reasoning Beyond Natural Language16 Jul 2024 0 repositories listed
-
Token-Supervised Value Models for Enhancing Mathematical Reasoning Capabilities of Large Language Models12 Jul 2024 0 repositories listed
-
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist11 Jul 2024 0 repositories listed
-
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On11 Jul 2024 0 repositories listed
-
When is the consistent prediction likely to be a correct prediction?8 Jul 2024 0 repositories listed
-
Question-Analysis Prompting Improves LLM Performance in Reasoning Tasks4 Jul 2024 0 repositories listed
-
AgentInstruct: Toward Generative Teaching with Agentic Flows3 Jul 2024 0 repositories listed
-
Turning Up the Heat: Min-p Sampling for Creative and Coherent LLM Outputs1 Jul 2024 0 repositories listed
-
Advancing Process Verification for Large Language Models via Tree-Based Preference Learning29 Jun 2024 0 repositories listed
-
LiteSearch: Efficacious Tree Search for LLM29 Jun 2024 0 repositories listed
-
PORT: Preference Optimization on Reasoning Traces23 Jun 2024 0 repositories listed
-
Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning20 Jun 2024 0 repositories listed
-
Uncertainty Aware Learning for Language Model Alignment7 Jun 2024 0 repositories listed
-
Does your data spark joy? Performance gains from domain upsampling at the end of training5 Jun 2024 0 repositories listed
-
Improve Mathematical Reasoning in Language Models by Automated Process Supervision5 Jun 2024 0 repositories listed
-
SpecDec++: Boosting Speculative Decoding via Adaptive Candidate Lengths30 May 2024 0 repositories listed
-
Arithmetic Reasoning with LLM: Prolog Generation & Permutation28 May 2024 0 repositories listed
-
Multi-Reference Preference Optimization for Large Language Models26 May 2024 0 repositories listed
-
MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time25 May 2024 0 repositories listed
-
Metacognitive Capabilities of LLMs: An Exploration in Mathematical Problem Solving20 May 2024 0 repositories listed
-
Meaning-Typed Programming: Language Abstraction and Runtime for Model-Integrated Applications14 May 2024 0 repositories listed
-
MathDivide: Improved mathematical reasoning by large language models12 May 2024 0 repositories listed
-
MAmmoTH2: Scaling Instructions from the Web6 May 2024 0 repositories listed
-
A Careful Examination of Large Language Model Performance on Grade School Arithmetic1 May 2024 0 repositories listed
-
Iterative Reasoning Preference Optimization30 Apr 2024 0 repositories listed
-
PARAMANU-GANITA: Language Model with Mathematical Capabilities22 Apr 2024 0 repositories listed
-
Relevant or Random: Can LLMs Truly Perform Analogical Reasoning?19 Apr 2024 0 repositories listed
-
Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models18 Apr 2024 0 repositories listed
-
Efficient Contextual LLM Cascades through Budget-Constrained Policy Learning17 Apr 2024 0 repositories listed
-
Automatic Prompt Selection for Large Language Models3 Apr 2024 0 repositories listed
-
Prompt-SAW: Leveraging Relation-Aware Graphs for Textual Prompt Compression30 Mar 2024 0 repositories listed
-
Supervisory Prompt Training26 Mar 2024 0 repositories listed
-
Self-Consistency Boosts Calibration for Math Reasoning14 Mar 2024 0 repositories listed
-
Prompt Selection and Augmentation for Few Examples Code Generation in Large Language Model and its Application in Robotics Control11 Mar 2024 0 repositories listed
-
4 Mar 2024 0 repositories listed
-
MathGenie: Generating Synthetic Data with Question Back-translation for Enhancing Mathematical Reasoning of LLMs26 Feb 2024 0 repositories listed
-
Look Before You Leap: Problem Elaboration Prompting Improves Mathematical Reasoning in Large Language Models24 Feb 2024 0 repositories listed
-
Fine-Grained Self-Endorsement Improves Factuality and Reasoning23 Feb 2024 0 repositories listed
-
SymBa: Symbolic Backward Chaining for Structured Natural Language Reasoning20 Feb 2024 0 repositories listed
-
Can Separators Improve Chain-of-Thought Prompting?16 Feb 2024 0 repositories listed
-
16 Feb 2024 0 repositories listed
-
Premise Order Matters in Reasoning with Large Language Models14 Feb 2024 0 repositories listed
-
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements13 Feb 2024 0 repositories listed
-
9 Feb 2024 0 repositories listed
-
RevOrder: A Novel Method for Enhanced Arithmetic in Language Models6 Feb 2024 0 repositories listed