Browse State-of-the-Art › Mathematical Reasoning › Papers, page 5
Mathematical Reasoning
Papers archive 2025-07-28
archive papers tagged: 805 · with a code link: 395 · where Syntology ran a sample: 197 (159 with a run with no instrument failure, 38 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (197 of 805 tagged: 159 with a run with no instrument failure, 38 where every run was a failure of Syntology's instrument)
Page 5 of 9: papers 401 to 500 of 805, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
CoRE: Enhancing Metacognition with Label-free Self-evaluation in LRMs8 Jul 2025 0 repositories listed
-
Large Language Models Don't Make Sense of Word Problems. A Scoping Review from a Mathematics Education Perspective30 Jun 2025 0 repositories listed
-
Layer Importance for Mathematical Reasoning is Forged in Pre-Training and Invariant after Post-Training27 Jun 2025 0 repositories listed
-
Inside you are many wolves: Using cognitive models to interpret value trade-offs in LLMs25 Jun 2025 0 repositories listed
-
Test-time Scaling Techniques in Theoretical Physics -- A Comparison of Methods on the TPBench Dataset25 Jun 2025 0 repositories listed
-
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model23 Jun 2025 0 repositories listed
-
PhysUniBench: An Undergraduate-Level Physics Reasoning Benchmark for Multimodal Models21 Jun 2025 0 repositories listed
-
Towards Advanced Mathematical Reasoning for LLMs via First-Order Logic Theorem Proving20 Jun 2025 0 repositories listed
-
Massive Supervised Fine-tuning Experiments Reveal How Data, Layer, and Training Factors Shape LLM Alignment Quality17 Jun 2025 0 repositories listed
-
Revisiting Chain-of-Thought Prompting: Zero-shot Can Be Stronger than Few-shot17 Jun 2025 0 repositories listed
-
A Technical Study into Small Reasoning Language Models16 Jun 2025 0 repositories listed
-
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles16 Jun 2025 0 repositories listed
-
Eliciting Reasoning in Language Models with Cognitive Tools13 Jun 2025 0 repositories listed
-
Investigating the Potential of Large Language Model-Based Router Multi-Agent Architectures for Foundation Design Automation: A Task Classification and Expert Selection Study13 Jun 2025 0 repositories listed
-
LearnAlign: Reasoning Data Selection for Reinforcement Learning in Large Language Models Based on Improved Gradient Alignment13 Jun 2025 0 repositories listed
-
Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning12 Jun 2025 0 repositories listed
-
PREMISE: Scalable and Strategic Prompt Optimization for Efficient Mathematical Reasoning in Large Models12 Jun 2025 0 repositories listed
-
Slimming Down LLMs Without Losing Their Minds12 Jun 2025 0 repositories listed
-
TeleMath: A Benchmark for Large Language Models in Telecom Mathematical Problem Solving12 Jun 2025 0 repositories listed
-
Large Language Models for Design Structure Matrix Optimization11 Jun 2025 0 repositories listed
-
Omni-DPO: A Dual-Perspective Paradigm for Dynamic Preference Learning of LLMs11 Jun 2025 0 repositories listed
-
Towards Efficient and Effective Alignment of Large Language Models11 Jun 2025 0 repositories listed
-
A Survey on Large Language Models for Mathematical Reasoning10 Jun 2025 0 repositories listed
-
Large Language Models Have Intrinsic Meta-Cognition, but Need a Good Lens10 Jun 2025 0 repositories listed
-
Temporalizing Confidence: Evaluation of Chain-of-Thought Reasoning with Signal Temporal Logic9 Jun 2025 0 repositories listed
-
Can Theoretical Physics Research Benefit from Language Agents?6 Jun 2025 0 repositories listed
-
Beyond Accuracy: Dissecting Mathematical Reasoning for LLMs Under Reinforcement Learning5 Jun 2025 0 repositories listed
-
Multi-Layer GRPO: Enhancing Reasoning and Self-Correction in Large Language Models5 Jun 2025 0 repositories listed
-
ProRefine: Inference-time Prompt Refinement with Textual Feedback5 Jun 2025 0 repositories listed
-
Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning5 Jun 2025 0 repositories listed
-
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos5 Jun 2025 0 repositories listed
-
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks2 Jun 2025 0 repositories listed
-
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking1 Jun 2025 0 repositories listed
-
1 Jun 2025 0 repositories listed Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Evaluation of LLMs for mathematical problem solving30 May 2025 0 repositories listed
-
AutoGPS: Automated Geometry Problem Solving via Multimodal Formalization and Deductive Reasoning29 May 2025 0 repositories listed
-
Diversity-Aware Policy Optimization for Large Language Model Reasoning29 May 2025 0 repositories listed
-
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability29 May 2025 0 repositories listed
-
Revisiting Multi-Agent Debate as Test-Time Scaling: A Systematic Study of Conditional Effectiveness29 May 2025 0 repositories listed
-
Revisiting Overthinking in Long Chain-of-Thought from the Perspective of Self-Doubt29 May 2025 0 repositories listed
-
Don't Think Longer, Think Wisely: Optimizing Thinking Dynamics for Large Reasoning Models27 May 2025 0 repositories listed
-
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles26 May 2025 0 repositories listed
-
HS-STAR: Hierarchical Sampling for Self-Taught Reasoners via Difficulty Estimation and Budget Reallocation26 May 2025 0 repositories listed
-
Improving Multilingual Math Reasoning for African Languages26 May 2025 0 repositories listed
-
ActiveDPO: Active Direct Preference Optimization for Sample-Efficient Alignment25 May 2025 0 repositories listed
-
AI4Math: A Native Spanish Benchmark for University-Level Mathematical Reasoning in Large Language Models25 May 2025 0 repositories listed
-
Don't Look Only Once: Towards Multimodal Interactive Reasoning with Selective Visual Revisitation24 May 2025 0 repositories listed
-
Efficient Long CoT Reasoning in Small Language Models24 May 2025 0 repositories listed
-
LogicCat: A Chain-of-Thought Text-to-SQL Benchmark for Multi-Domain Reasoning Challenges24 May 2025 0 repositories listed
-
Guided by Gut: Efficient Test-Time Scaling with Reinforced Intrinsic Confidence23 May 2025 0 repositories listed
-
PPT: A Process-based Preference Learning Framework for Self Improving Table Question Answering Models23 May 2025 0 repositories listed
-
The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs23 May 2025 0 repositories listed
-
Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN22 May 2025 0 repositories listed
-
Bottlenecked Transformers: Periodic KV Cache Abstraction for Generalised Reasoning22 May 2025 0 repositories listed
-
Dynamic Sampling that Adapts: Iterative DPO for Self-Aware Mathematical Reasoning22 May 2025 0 repositories listed
-
HOFT: Householder Orthogonal Fine-tuning22 May 2025 0 repositories listed
-
MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models22 May 2025 0 repositories listed
-
SMART: Self-Generating and Self-Validating Multi-Dimensional Assessment for LLMs' Mathematical Problem Solving22 May 2025 0 repositories listed
-
Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains22 May 2025 0 repositories listed
-
Can LLMs understand Math? -- Exploring the Pitfalls in Mathematical Reasoning21 May 2025 0 repositories listed
-
Learning to Rank Chain-of-Thought: An Energy-Based Approach with Outcome Supervision21 May 2025 0 repositories listed
-
MAPS: A Multilingual Benchmark for Global Agent Performance and Security21 May 2025 0 repositories listed
-
SSR: Speculative Parallel Scaling Reasoning in Test-time21 May 2025 0 repositories listed
-
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems21 May 2025 0 repositories listed
-
AAPO: Enhance the Reasoning Capabilities of LLMs with Advantage Momentum20 May 2025 0 repositories listed
-
Beyond the First Error: Process Reward Models for Reflective Mathematical Reasoning20 May 2025 0 repositories listed
-
Can Pruning Improve Reasoning? Revisiting Long-CoT Compression with Capability in Mind for Better Reasoning20 May 2025 0 repositories listed
-
DRP: Distilled Reasoning Pruning with Skill-aware Step Decomposition for Efficient Large Reasoning Models20 May 2025 0 repositories listed
-
Mind the Gap: Bridging Thought Leap for Improved Chain-of-Thought Tuning20 May 2025 0 repositories listed
-
OSoRA: Output-Dimension and Singular-Value Initialized Low-Rank Adaptation20 May 2025 0 repositories listed
-
Text Generation Beyond Discrete Token Sampling20 May 2025 0 repositories listed
-
WirelessMathBench: A Mathematical Modeling Benchmark for LLMs in Wireless Communications20 May 2025 0 repositories listed
-
AutoMathKG: The automated mathematical knowledge graph based on LLM and vector database19 May 2025 0 repositories listed
-
Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers19 May 2025 0 repositories listed
-
Guided Search Strategies in Non-Serializable Environments with Applications to Software Engineering Agents19 May 2025 0 repositories listed
-
Selective Code Generation for Functional Guarantees19 May 2025 0 repositories listed
-
Step-wise Adaptive Integration of Supervised Fine-tuning and Reinforcement Learning for Task-Specific LLMs19 May 2025 0 repositories listed
-
Unlocking the Potential of Difficulty Prior in RL-based Multimodal Reasoning19 May 2025 0 repositories listed
-
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization18 May 2025 0 repositories listed
-
Real-Time Verification of Embodied Reasoning for Generative Skill Acquisition16 May 2025 0 repositories listed
-
Are Large Language Models Robust in Understanding Code Against Semantics-Preserving Mutations?15 May 2025 0 repositories listed
-
Agent-as-a-Service based on Agent Network13 May 2025 0 repositories listed
-
Learning Like Humans: Advancing LLM Reasoning Capabilities via Adaptive Difficulty Curriculum Learning and Expert-Guided Self-Reformulation13 May 2025 0 repositories listed
-
Assessing Robustness to Spurious Correlations in Post-Training Language Models9 May 2025 0 repositories listed
-
Knowledge Augmented Complex Problem Solving with Large Language Models: A Survey6 May 2025 0 repositories listed
-
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think29 Apr 2025 0 repositories listed
-
RV-Syn: Rational and Verifiable Mathematical Reasoning Data Synthesis based on Structured Function Library29 Apr 2025 0 repositories listed
-
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets28 Apr 2025 0 repositories listed
-
Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning28 Apr 2025 0 repositories listed
-
SPC: Evolving Self-Play Critic via Adversarial Games for LLM Reasoning27 Apr 2025 0 repositories listed
-
25 Apr 2025 0 repositories listed Syntology 1 ran (of which 0 constructed an object rather than computing a result; 0 with no instrument failure: 0 honoured, 0 violated, 0 with no contract checked; 1 where Syntology's instrument failed) · 0 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
DeepDistill: Enhancing LLM Reasoning Capabilities via Large-Scale Difficulty-Graded Data Training24 Apr 2025 0 repositories listed
-
Evaluating Grounded Reasoning by Code-Assisted Large Language Models for Mathematics24 Apr 2025 0 repositories listed
-
Parameter-Efficient Checkpoint Merging via Metrics-Weighted Averaging23 Apr 2025 0 repositories listed
-
Improving RL Exploration for LLM Reasoning through Retrospective Replay19 Apr 2025 0 repositories listed
-
BitNet b1.58 2B4T Technical Report16 Apr 2025 0 repositories listed
-
Assessment of Evolving Large Language Models in Upper Secondary Mathematics15 Apr 2025 0 repositories listed
Syntology lines on 5 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.