Browse State-of-the-Art › Mathematical Reasoning › Papers, page 6
Mathematical Reasoning
Papers archive 2025-07-28
archive papers tagged: 805 · with a code link: 395 · where Syntology ran a sample: 197 (159 with a run with no instrument failure, 38 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (197 of 805 tagged: 159 with a run with no instrument failure, 38 where every run was a failure of Syntology's instrument)
Page 6 of 9: papers 501 to 600 of 805, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Enhancing Mathematical Reasoning in Large Language Models with Self-Consistency-Based Hallucination Detection13 Apr 2025 0 repositories listed
-
Supervised Optimism Correction: Be Confident When LLMs Are Sure10 Apr 2025 0 repositories listed
-
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use7 Apr 2025 0 repositories listed
-
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation4 Apr 2025 0 repositories listed
-
Sample, Don't Search: Rethinking Test-Time Alignment for Language Models4 Apr 2025 0 repositories listed
-
LexPam: Legal Procedure Awareness-Guided Mathematical Reasoning3 Apr 2025 0 repositories listed
-
LLM for Complex Reasoning Task: An Exploratory Study in Fermi Problems3 Apr 2025 0 repositories listed
-
LLM Library Learning Fails: A LEGO-Prover Case Study3 Apr 2025 0 repositories listed
-
Brains vs. Bytes: Evaluating LLM Proficiency in Olympiad Mathematics1 Apr 2025 0 repositories listed
-
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning1 Apr 2025 0 repositories listed
-
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study1 Apr 2025 0 repositories listed
-
VerifiAgent: a Unified Verification Agent in Language Model Reasoning1 Apr 2025 0 repositories listed
-
Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains31 Mar 2025 0 repositories listed
-
The Axiom-Based Atlas: A Structural Mapping of Theorems via Foundational Proof Vectors31 Mar 2025 0 repositories listed
-
Entropy-Aware Branching for Improved Mathematical Reasoning27 Mar 2025 0 repositories listed
-
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad27 Mar 2025 0 repositories listed
-
MATHGLANCE: Multimodal Large Language Models Do Not Know Where to Look in Mathematical Diagrams26 Mar 2025 0 repositories listed
-
Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking25 Mar 2025 0 repositories listed
-
Learning to chain-of-thought with Jensen's evidence lower bound25 Mar 2025 0 repositories listed
-
Process or Result? Manipulated Ending Tokens Can Mislead Reasoning LLMs to Ignore the Correct Reasoning Steps25 Mar 2025 0 repositories listed
-
RL-finetuning LLMs from on- and off-policy data with a single algorithm25 Mar 2025 0 repositories listed
-
CLEAR: Contrasting Textual Feedback with Experts and Amateurs for Reasoning24 Mar 2025 0 repositories listed
-
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning17 Mar 2025 0 repositories listed
-
Pensez: Less Data, Better Reasoning -- Rethinking French LLM17 Mar 2025 0 repositories listed
-
Reliable and Efficient Amortized Model-based Evaluation17 Mar 2025 0 repositories listed
-
Evaluating Mathematical Reasoning Across Large Language Models: A Fine-Grained Approach13 Mar 2025 0 repositories listed
-
7 Mar 2025 0 repositories listed Syntology 7 ran (of which 1 constructed an object rather than computing a result; 5 with no instrument failure: 1 honoured, 3 violated, 1 with no contract checked; 2 where Syntology's instrument failed) · 0 unverified (of 7 harvested samples) · 7 pointer-only (licence)
-
Speculative Decoding for Multi-Sample Inference7 Mar 2025 0 repositories listed
-
Better Process Supervision with Bi-directional Rewarding Signals6 Mar 2025 0 repositories listed
-
Towards Understanding Multi-Round Large Language Model Reasoning: Approximability, Learnability and Generalizability5 Mar 2025 0 repositories listed
-
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models4 Mar 2025 0 repositories listed
-
None of the Above, Less of the Right: Parallel Patterns between Humans and LLMs on Multi-Choice Questions Answering3 Mar 2025 0 repositories listed
-
MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts28 Feb 2025 0 repositories listed
-
Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models27 Feb 2025 0 repositories listed
-
Multi2: Multi-Agent Test-Time Scalable Framework for Multi-Document Processing27 Feb 2025 0 repositories listed
-
Revisiting Self-Consistency from Dynamic Distributional Alignment Perspective on Answer Aggregation27 Feb 2025 0 repositories listed
-
Thinking Slow, Fast: Scaling Inference Compute with Distilled Reasoners27 Feb 2025 0 repositories listed
-
Weaker LLMs' Opinions Also Matter: Mixture of Opinions Enhances LLM's Mathematical Reasoning26 Feb 2025 0 repositories listed
-
LeanProgress: Guiding Search for Neural Theorem Proving via Proof Progress Prediction25 Feb 2025 0 repositories listed
-
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning25 Feb 2025 0 repositories listed
-
Full-Step-DPO: Self-Supervised Preference Optimization with Step-wise Rewards for Mathematical Reasoning20 Feb 2025 0 repositories listed
-
Retrieval-Augmented Process Reward Model for Generalizable Mathematical Reasoning20 Feb 2025 0 repositories listed
-
From Correctness to Comprehension: AI Agents for Personalized Error Diagnosis in Education19 Feb 2025 0 repositories listed
-
Integrating Arithmetic Learning Improves Mathematical Reasoning in Smaller Models18 Feb 2025 0 repositories listed
-
Sens-Merging: Sensitivity-Guided Parameter Balancing for Merging Large Language Models18 Feb 2025 0 repositories listed
-
Theorem Prover as a Judge for Synthetic Data Generation18 Feb 2025 0 repositories listed
-
Large Language Models and Mathematical Reasoning Failures17 Feb 2025 0 repositories listed
-
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task17 Feb 2025 0 repositories listed
-
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving17 Feb 2025 0 repositories listed
-
Leveraging Constrained Monte Carlo Tree Search to Generate Reliable Long Chain-of-Thought for Mathematical Reasoning16 Feb 2025 0 repositories listed
-
Uncertainty-Aware Step-wise Verification with Generative Reward Models16 Feb 2025 0 repositories listed
-
1bit-Merging: Dynamic Quantized Merging for Large Language Models15 Feb 2025 0 repositories listed
-
Evaluating the Meta- and Object-Level Reasoning of Large Language Models for Question Answering14 Feb 2025 0 repositories listed
-
GoRA: Gradient-driven Adaptive Low Rank Adaptation13 Feb 2025 0 repositories listed
-
LLMs can implicitly learn from mistakes in-context12 Feb 2025 0 repositories listed
-
One Example Shown, Many Concepts Known! Counterexample-Driven Conceptual Reasoning in Mathematical LLMs12 Feb 2025 0 repositories listed
-
Selective Self-to-Supervised Fine-Tuning for Generalization in Large Language Models12 Feb 2025 0 repositories listed
-
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations10 Feb 2025 0 repositories listed
-
Self-Training Large Language Models for Tool-Use Without Demonstrations9 Feb 2025 0 repositories listed
-
Evolving LLMs' Self-Refinement Capability via Iterative Preference Optimization8 Feb 2025 0 repositories listed
-
LLMs can be easily Confused by Instructional Distractions5 Feb 2025 0 repositories listed
-
Path Planning for Masked Diffusion Model Sampling5 Feb 2025 0 repositories listed
-
Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment5 Feb 2025 0 repositories listed
-
Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning5 Feb 2025 0 repositories listed
-
Policy Guided Tree Search for Enhanced LLM Reasoning4 Feb 2025 0 repositories listed
-
Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMs4 Feb 2025 0 repositories listed
-
Satori: Reinforcement Learning with Chain-of-Action-Thought Enhances LLM Reasoning via Autoregressive Search4 Feb 2025 0 repositories listed
-
MergeME: Model Merging Techniques for Homogeneous and Heterogeneous MoEs3 Feb 2025 0 repositories listed
-
Language Models Use Trigonometry to Do Addition2 Feb 2025 0 repositories listed
-
Improving Rule-based Reasoning in LLMs via Neurosymbolic Representations31 Jan 2025 0 repositories listed
-
From Informal to Formal -- Incorporating and Evaluating LLMs on Natural Language Requirements to Verifiable Formal Proofs27 Jan 2025 0 repositories listed
-
LemmaHead: RAG Assisted Proof Generation Using Large Language Models27 Jan 2025 0 repositories listed
-
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework26 Jan 2025 0 repositories listed
-
The Karp Dataset24 Jan 2025 0 repositories listed
-
Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages23 Jan 2025 0 repositories listed
-
Coarse-to-Fine Process Reward Modeling for Enhanced Mathematical Reasoning23 Jan 2025 0 repositories listed
-
23 Jan 2025 0 repositories listed Syntology 0 ran · 2 unverified (of 2 harvested samples) · 2 pointer-only (licence)
-
CDW-CoT: Clustered Distance-Weighted Chain-of-Thoughts Reasoning21 Jan 2025 0 repositories listed
-
Benchmarking Large Language Models via Random Variables20 Jan 2025 0 repositories listed
-
Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective19 Jan 2025 0 repositories listed
-
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback18 Jan 2025 0 repositories listed
-
The Lessons of Developing Process Reward Models in Mathematical Reasoning13 Jan 2025 0 repositories listed
-
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning6 Jan 2025 0 repositories listed
-
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap5 Jan 2025 0 repositories listed
-
Table as Thought: Exploring Structured Thoughts in LLM Reasoning4 Jan 2025 0 repositories listed
-
Enhancing Reasoning through Process Supervision with Monte Carlo Tree Search2 Jan 2025 0 repositories listed
-
Plug-and-Play Training Framework for Preference Optimization30 Dec 2024 0 repositories listed
-
LLM Reasoning Engine: Specialized Training for Enhanced Mathematical Reasoning28 Dec 2024 0 repositories listed
-
System-2 Mathematical Reasoning via Enriched Instruction Tuning22 Dec 2024 0 repositories listed
-
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning20 Dec 2024 0 repositories listed
-
Formal Mathematical Reasoning: A New Frontier in AI20 Dec 2024 0 repositories listed
-
What Are Step-Level Reward Models Rewarding? Counterintuitive Findings from MCTS-Boosted Mathematical Reasoning20 Dec 2024 0 repositories listed
-
Channel Merging: Preserving Specialization for Merged Experts18 Dec 2024 0 repositories listed
-
MetaRuleGPT: Recursive Numerical Reasoning of Language Models Trained with Simple Rules18 Dec 2024 0 repositories listed
-
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges16 Dec 2024 0 repositories listed
-
Can Language Models Rival Mathematics Students? Evaluating Mathematical Reasoning through Textual Manipulation and Human Experiments16 Dec 2024 0 repositories listed
-
Low-Rank Adaptation with Task-Relevant Feature Enhancement for Fine-tuning Language Models13 Dec 2024 0 repositories listed
-
A Graph-Based Synthetic Data Pipeline for Scaling High-Quality Reasoning Instructions12 Dec 2024 0 repositories listed
-
Sail into the Headwind: Alignment via Robust Rewards and Dynamic Labels against Reward Hacking12 Dec 2024 0 repositories listed
-
SmolTulu: Higher Learning Rate to Batch Size Ratios Can Lead to Better Reasoning in SLMs11 Dec 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.