Browse State-of-the-Art › GSM8K › Papers, page 3
GSM8K
Papers archive 2025-07-28
archive papers tagged: 439 · with a code link: 209 · where Syntology ran a sample: 116 (96 with a run with no instrument failure, 20 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (116 of 439 tagged: 96 with a run with no instrument failure, 20 where every run was a failure of Syntology's instrument)
Page 3 of 5: papers 201 to 300 of 439, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
23 May 2023 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
23 May 2023 1 repository listed
-
23 May 2023 1 repository listed
-
19 Apr 2023 1 repository listed Syntology official (archive's flag): 3 ran · 3 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 4 pointer-only (licence)
-
16 Apr 2023 1 repository listed
-
12 Apr 2023 1 repository listed
-
27 Jan 2023 1 repository listed Syntology official (archive's flag): 1 ran · 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 1 unverified (of 2 harvested samples)
-
1 Dec 2022 1 repository listed
-
28 May 2022 1 repository listed Syntology official (archive's flag): 7 ran · 7 ran (of which 0 constructed an object rather than computing a result; 7 with no instrument failure: 0 honoured, 0 violated, 7 with no contract checked; 0 where Syntology's instrument failed) · 4 unverified (of 11 harvested samples)
-
GEMMAS: Graph-based Evaluation Metrics for Multi Agent Systems17 Jul 2025 0 repositories listed
-
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?15 Jul 2025 0 repositories listed
-
CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs8 Jul 2025 0 repositories listed
-
Activation Steering for Chain-of-Thought Compression7 Jul 2025 0 repositories listed
-
Scaling Speculative Decoding with Lookahead Reasoning24 Jun 2025 0 repositories listed
-
23 Jun 2025 0 repositories listed Syntology 4 ran (of which 0 constructed an object rather than computing a result; 1 with no instrument failure: 0 honoured, 1 violated, 0 with no contract checked; 3 where Syntology's instrument failed) · 1 unverified (of 5 harvested samples) · 4 pointer-only (licence)
-
Fractional Reasoning via Latent Steering Vectors Improves Inference Time Compute18 Jun 2025 0 repositories listed
-
Excessive Reasoning Attack on Reasoning LLMs17 Jun 2025 0 repositories listed
-
LoRA-Mixer: Coordinate Modular LoRA Experts Through Serial Attention Routing17 Jun 2025 0 repositories listed
-
LearnAlign: Reasoning Data Selection for Reinforcement Learning in Large Language Models Based on Improved Gradient Alignment13 Jun 2025 0 repositories listed
-
Fast on the Easy, Deep on the Hard: Efficient Reasoning via Powered Length Penalty12 Jun 2025 0 repositories listed
-
PREMISE: Scalable and Strategic Prompt Optimization for Efficient Mathematical Reasoning in Large Models12 Jun 2025 0 repositories listed
-
Slimming Down LLMs Without Losing Their Minds12 Jun 2025 0 repositories listed
-
Enhancing Reasoning Capabilities of Small Language Models with Blueprints and Prompt Template Search10 Jun 2025 0 repositories listed
-
Guideline Forest: Experience-Induced Multi-Guideline Reasoning with Stepwise Aggregation9 Jun 2025 0 repositories listed
-
Text-to-LoRA: Instant Transformer Adaption6 Jun 2025 0 repositories listed
-
Automatic Robustness Stress Testing of LLMs as Mathematical Problem Solvers5 Jun 2025 0 repositories listed
-
Evaluation of LLMs for mathematical problem solving30 May 2025 0 repositories listed
-
Model Unlearning via Sparse Autoencoder Subspace Guided Projections30 May 2025 0 repositories listed
-
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation29 May 2025 0 repositories listed
-
CoThink: Token-Efficient Reasoning via Instruct Models Guiding Reasoning Models28 May 2025 0 repositories listed
-
Maximizing Confidence Alone Improves Reasoning28 May 2025 0 repositories listed
-
Efficient Data Selection at Scale via Influence Distillation25 May 2025 0 repositories listed
-
LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models25 May 2025 0 repositories listed
-
System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic Shortcuts25 May 2025 0 repositories listed
-
Steering LLM Reasoning Through Bias-Only Adaptation24 May 2025 0 repositories listed
-
PMPO: Probabilistic Metric Prompt Optimization for Small and Large Language Models22 May 2025 0 repositories listed
-
Learning to Rank Chain-of-Thought: An Energy-Based Approach with Outcome Supervision21 May 2025 0 repositories listed
-
DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models20 May 2025 0 repositories listed
-
Dual Decomposition of Weights and Singular Value Low Rank Adaptation20 May 2025 0 repositories listed
-
Self-Reasoning Language Models: Unfold Hidden Reasoning Chains with Few Reasoning Catalyst20 May 2025 0 repositories listed
-
RL in Name Only? Analyzing the Structural Assumptions in RL post-training for LLMs19 May 2025 0 repositories listed
-
Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language Models15 May 2025 0 repositories listed
-
Accelerating Chain-of-Thought Reasoning: When Goal-Gradient Importance Meets Dynamic Skipping13 May 2025 0 repositories listed
-
AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection12 May 2025 0 repositories listed
-
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models12 May 2025 0 repositories listed
-
Elastic Weight Consolidation for Full-Parameter Continual Pre-Training of Gemma29 May 2025 0 repositories listed
-
Memory-Efficient LLM Training by Various-Grained Low-Rank Projection of Gradients3 May 2025 0 repositories listed
-
Efficient Fine-Tuning of Quantized Models via Adaptive Rank and Bitwidth2 May 2025 0 repositories listed
-
Local Prompt Optimization29 Apr 2025 0 repositories listed
-
Trace-of-Thought Prompting: Investigating Prompt-Based Knowledge Distillation Through Question Decomposition29 Apr 2025 0 repositories listed
-
AutoJudge: Judge Decoding Without Manual Annotation28 Apr 2025 0 repositories listed
-
Training Large Language Models to Reason via EM Policy Gradient24 Apr 2025 0 repositories listed
-
Not All Rollouts are Useful: Down-Sampling Rollouts in LLM Reinforcement Learning18 Apr 2025 0 repositories listed
-
Entropy-Guided Watermarking for LLMs: A Test-Time Framework for Robust and Traceable Text Generation16 Apr 2025 0 repositories listed
-
Question Tokens Deserve More Attention: Enhancing Large Language Models without Training through Step-by-Step Reading and Question Attention Recalibration13 Apr 2025 0 repositories listed
-
Supervised Optimism Correction: Be Confident When LLMs Are Sure10 Apr 2025 0 repositories listed
-
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use7 Apr 2025 0 repositories listed
-
Sample, Don't Search: Rethinking Test-Time Alignment for Language Models4 Apr 2025 0 repositories listed
-
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency4 Apr 2025 0 repositories listed
-
D²LoRA: Data-Driven LoRA Initialization for Low Resource Tasks23 Mar 2025 0 repositories listed
-
Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs18 Mar 2025 0 repositories listed
-
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach17 Mar 2025 0 repositories listed
-
Rule-Guided Feedback: Enhancing Reasoning by Enforcing Rule Adherence in Large Language Models14 Mar 2025 0 repositories listed
-
Position-Aware Depth Decay Decoding (D³): Boosting Large Language Model Inference Efficiency11 Mar 2025 0 repositories listed
-
SOLAR: Scalable Optimization of Large-scale Architecture for Reasoning6 Mar 2025 0 repositories listed
-
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models4 Mar 2025 0 repositories listed
-
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge27 Feb 2025 0 repositories listed
-
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?26 Feb 2025 0 repositories listed
-
Weaker LLMs' Opinions Also Matter: Mixture of Opinions Enhances LLM's Mathematical Reasoning26 Feb 2025 0 repositories listed
-
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models25 Feb 2025 0 repositories listed
-
LED-Merging: Mitigating Safety-Utility Conflicts in Model Merging with Location-Election-Disjoint24 Feb 2025 0 repositories listed
-
Dynamic Parallel Tree Search for Efficient LLM Reasoning22 Feb 2025 0 repositories listed
-
From Correctness to Comprehension: AI Agents for Personalized Error Diagnosis in Education19 Feb 2025 0 repositories listed
-
Integrating Arithmetic Learning Improves Mathematical Reasoning in Smaller Models18 Feb 2025 0 repositories listed
-
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task17 Feb 2025 0 repositories listed
-
Balancing the Budget: Understanding Trade-offs Between Supervised and Preference-Based Finetuning16 Feb 2025 0 repositories listed
-
Leveraging Uncertainty Estimation for Efficient LLM Routing16 Feb 2025 0 repositories listed
-
Uncertainty-Aware Search and Value Models: Mitigating Search Scaling Flaws in LLMs16 Feb 2025 0 repositories listed
-
Hybrid Offline-online Scheduling Method for Large Language Model Inference Optimization14 Feb 2025 0 repositories listed
-
Cost-Saving LLM Cascades with Early Abstention13 Feb 2025 0 repositories listed
-
Self-Training Large Language Models for Tool-Use Without Demonstrations9 Feb 2025 0 repositories listed
-
Evolving LLMs' Self-Refinement Capability via Iterative Preference Optimization8 Feb 2025 0 repositories listed
-
Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment5 Feb 2025 0 repositories listed
-
BARE: Leveraging Base Language Models for Few-Shot Synthetic Data Generation3 Feb 2025 0 repositories listed
-
ChunkKV: Semantic-Preserving KV Cache Compression for Efficient Long-Context LLM Inference1 Feb 2025 0 repositories listed
-
Pheromone-based Learning of Optimal Reasoning Paths31 Jan 2025 0 repositories listed
-
RotateKV: Accurate and Robust 2-Bit KV Cache Quantization for LLMs via Outlier-Aware Adaptive Rotations25 Jan 2025 0 repositories listed
-
Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs21 Jan 2025 0 repositories listed
-
DNA 1.0 Technical Report18 Jan 2025 0 repositories listed
-
Semantic Exploration with Adaptive Gating for Efficient Problem Solving with Language Models10 Jan 2025 0 repositories listed
-
InfiFusion: A Unified Framework for Enhanced Cross-Model Reasoning via LLM Fusion6 Jan 2025 0 repositories listed
-
Recursive Decomposition of Logical Thoughts: Framework for Superior Reasoning and Knowledge Propagation in Large Language Models3 Jan 2025 0 repositories listed
-
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs30 Dec 2024 0 repositories listed
-
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning23 Dec 2024 0 repositories listed
-
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions22 Dec 2024 0 repositories listed
-
System-2 Mathematical Reasoning via Enriched Instruction Tuning22 Dec 2024 0 repositories listed
-
Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree17 Dec 2024 0 repositories listed
-
A Graph-Based Synthetic Data Pipeline for Scaling High-Quality Reasoning Instructions12 Dec 2024 0 repositories listed
-
Learning to Reason via Self-Iterative Process Feedback for Small Language Models11 Dec 2024 0 repositories listed
-
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs11 Dec 2024 0 repositories listed
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.