Browse State-of-the-Art › Math › Papers, page 9
Math
Papers archive 2025-07-28
archive papers tagged: 1,596 · with a code link: 765 · where Syntology ran a sample: 349 (286 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (349 of 1,596 tagged: 286 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument)
Page 9 of 16: papers 801 to 900 of 1,596, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Automatic Robustness Stress Testing of LLMs as Mathematical Problem Solvers5 Jun 2025 0 repositories listed
-
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models5 Jun 2025 0 repositories listed
-
Perceptual Decoupling for Scalable Multi-modal Reasoning via Reward-Optimized Captioning5 Jun 2025 0 repositories listed
-
Simulating LLM-to-LLM Tutoring for Multilingual Math Feedback5 Jun 2025 0 repositories listed
-
TreeRPO: Tree Relative Policy Optimization5 Jun 2025 0 repositories listed
-
Rectified Sparse Attention4 Jun 2025 0 repositories listed
-
MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching3 Jun 2025 0 repositories listed
-
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem3 Jun 2025 0 repositories listed
-
Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains2 Jun 2025 0 repositories listed
-
SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis2 Jun 2025 0 repositories listed
-
GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking1 Jun 2025 0 repositories listed
-
Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking30 May 2025 0 repositories listed
-
Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning30 May 2025 0 repositories listed
-
Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic Computation29 May 2025 0 repositories listed
-
DINGO: Constrained Inference for Diffusion LLMs29 May 2025 0 repositories listed
-
Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models29 May 2025 0 repositories listed
-
Let's Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM's Math Capability29 May 2025 0 repositories listed
-
Matryoshka Model Learning for Improved Elastic Student Models29 May 2025 0 repositories listed
-
PBEBench: A Multi-Step Programming by Examples Reasoning Benchmark inspired by Historical Linguistics29 May 2025 0 repositories listed
-
Maximizing Confidence Alone Improves Reasoning28 May 2025 0 repositories listed
-
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning27 May 2025 0 repositories listed
-
Done Is Better than Perfect: Unlocking Efficient Reasoning by Structured Multi-Turn Decomposition26 May 2025 0 repositories listed
-
Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles26 May 2025 0 repositories listed
-
Faster and Better LLMs via Latency-Aware Test-Time Scaling26 May 2025 0 repositories listed
-
Improving Multilingual Math Reasoning for African Languages26 May 2025 0 repositories listed
-
Interleaved Reasoning for Large Language Models via Reinforcement Learning26 May 2025 0 repositories listed
-
Prismatic Synthesis: Gradient-based Data Diversification Boosts Generalization in LLM Reasoning26 May 2025 0 repositories listed
-
The Role of Diversity in In-Context Learning for Large Language Models26 May 2025 0 repositories listed
-
Which Data Attributes Stimulate Math and Code Reasoning? An Investigation via Influence Functions26 May 2025 0 repositories listed
-
AI4Math: A Native Spanish Benchmark for University-Level Mathematical Reasoning in Large Language Models25 May 2025 0 repositories listed
-
Anchored Diffusion Language Model24 May 2025 0 repositories listed
-
24 May 2025 0 repositories listed Syntology 3 ran (of which 0 constructed an object rather than computing a result; 3 with no instrument failure: 0 honoured, 0 violated, 3 with no contract checked; 0 where Syntology's instrument failed) · 1 unverified (of 4 harvested samples) · 1 pointer-only (licence)
-
MSA at BEA 2025 Shared Task: Disagreement-Aware Instruction Tuning for Multi-Dimensional Evaluation of LLMs as Math Tutors24 May 2025 0 repositories listed
-
On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization24 May 2025 0 repositories listed
-
Steering LLM Reasoning Through Bias-Only Adaptation24 May 2025 0 repositories listed
-
More Thinking, Less Seeing? Assessing Amplified Hallucination in Multimodal Reasoning Models23 May 2025 0 repositories listed
-
One RL to See Them All: Visual Triple Unified Reinforcement Learning23 May 2025 0 repositories listed
-
Outcome-based Reinforcement Learning to Predict the Future23 May 2025 0 repositories listed
-
The Unreasonable Effectiveness of Model Merging for Cross-Lingual Transfer in LLMs23 May 2025 0 repositories listed
-
VideoGameBench: Can Vision-Language Models complete popular video games?23 May 2025 0 repositories listed
-
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning22 May 2025 0 repositories listed
-
Incremental Sequence Classification with Temporal Consistency22 May 2025 0 repositories listed
-
RBench-V: A Primary Assessment for Visual Reasoning Models with Multi-modal Outputs22 May 2025 0 repositories listed
-
Veracity Bias and Beyond: Uncovering LLMs' Hidden Beliefs in Problem-Solving Reasoning22 May 2025 0 repositories listed
-
Can LLMs understand Math? -- Exploring the Pitfalls in Mathematical Reasoning21 May 2025 0 repositories listed
-
Learning to Rank Chain-of-Thought: An Energy-Based Approach with Outcome Supervision21 May 2025 0 repositories listed
-
MAPS: A Multilingual Benchmark for Global Agent Performance and Security21 May 2025 0 repositories listed
-
SSR: Speculative Parallel Scaling Reasoning in Test-time21 May 2025 0 repositories listed
-
Thought-Augmented Policy Optimization: Bridging External Guidance and Internal Capabilities21 May 2025 0 repositories listed
-
Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems21 May 2025 0 repositories listed
-
EasyMath: A 0-shot Math Benchmark for SLMs20 May 2025 0 repositories listed
-
RL of Thoughts: Navigating LLM Reasoning with Inference-time Reinforcement Learning20 May 2025 0 repositories listed
-
The Hallucination Tax of Reinforcement Finetuning20 May 2025 0 repositories listed
-
Unearthing Gems from Stones: Policy Optimization with Negative Sample Augmentation for LLM Reasoning20 May 2025 0 repositories listed
-
AutoMathKG: The automated mathematical knowledge graph based on LLM and vector database19 May 2025 0 repositories listed
-
SEED-GRPO: Semantic Entropy Enhanced GRPO for Uncertainty-Aware Policy Optimization18 May 2025 0 repositories listed
-
LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades17 May 2025 0 repositories listed
-
MoL for LLMs: Dual-Loss Optimization to Enhance Domain Expertise While Preserving General Capabilities17 May 2025 0 repositories listed
-
SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning16 May 2025 0 repositories listed
-
Critique-Guided Distillation: Improving Supervised Fine-tuning via Better Distillation16 May 2025 0 repositories listed
-
DIF: A Framework for Benchmarking and Verifying Implicit Bias in LLMs15 May 2025 0 repositories listed
-
Reinforcing the Diffusion Chain of Lateral Thought with Diffusion Language Models15 May 2025 0 repositories listed
-
Accelerating Chain-of-Thought Reasoning: When Goal-Gradient Importance Meets Dynamic Skipping13 May 2025 0 repositories listed
-
Measurement to Meaning: A Validity-Centered Framework for AI Evaluation13 May 2025 0 repositories listed
-
Learning from Peers in Reasoning Models12 May 2025 0 repositories listed
-
Multimodal Assessment of Classroom Discourse Quality: A Text-Centered Attention-Based Multi-Task Learning Approach12 May 2025 0 repositories listed
-
S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models12 May 2025 0 repositories listed
-
DialogueReason: Rule-Based RL Sparks Dialogue Reasoning in LLMs11 May 2025 0 repositories listed
-
xGen-small Technical Report10 May 2025 0 repositories listed
-
Generative Discovery of Partial Differential Equations by Learning from Math Handbooks9 May 2025 0 repositories listed
-
Scalable LLM Math Reasoning Acceleration with Low-rank Distillation8 May 2025 0 repositories listed
-
Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers7 May 2025 0 repositories listed
-
A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law5 May 2025 0 repositories listed
-
Generating Narrated Lecture Videos from Slides with Synchronized Highlights5 May 2025 0 repositories listed
-
SIMPLEMIX: Frustratingly Simple Mixing of Off- and On-policy Data in Language Model Preference Learning5 May 2025 0 repositories listed
-
LookAlike: Consistent Distractor Generation in Math MCQs3 May 2025 0 repositories listed
-
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models30 Apr 2025 0 repositories listed
-
LLMs Do Not Have Human-Like Working Memory30 Apr 2025 0 repositories listed
-
Phi-4-Mini-Reasoning: Exploring the Limits of Small Reasoning Language Models in Math30 Apr 2025 0 repositories listed
-
Phi-4-reasoning Technical Report30 Apr 2025 0 repositories listed
-
Local Prompt Optimization29 Apr 2025 0 repositories listed
-
Trace-of-Thought Prompting: Investigating Prompt-Based Knowledge Distillation Through Question Decomposition29 Apr 2025 0 repositories listed
-
Accurate and Diverse LLM Mathematical Reasoning via Automated PRM-Guided GFlowNets28 Apr 2025 0 repositories listed
-
Evaluating Grounded Reasoning by Code-Assisted Large Language Models for Mathematics24 Apr 2025 0 repositories listed
-
Training Large Language Models to Reason via EM Policy Gradient24 Apr 2025 0 repositories listed
-
SplitReason: Learning To Offload Reasoning23 Apr 2025 0 repositories listed
-
LongPerceptualThoughts: Distilling System-2 Reasoning for System-1 Perception21 Apr 2025 0 repositories listed
-
OTC: Optimal Tool Calls via Reinforcement Learning21 Apr 2025 0 repositories listed
-
Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?18 Apr 2025 0 repositories listed
-
Enhancing Math Learning in an LMS Using AI-Driven Question Recommendations18 Apr 2025 0 repositories listed
-
In between myth and reality: AI for math -- a case study in category theory17 Apr 2025 0 repositories listed
-
MathPhys-Guided Coarse-to-Fine Anomaly Synthesis with SQE-Driven Bi-Level Optimization for Anomaly Detection17 Apr 2025 0 repositories listed
-
THOUGHTTERMINATOR: Benchmarking, Calibrating, and Mitigating Overthinking in Reasoning Models17 Apr 2025 0 repositories listed
-
Entropy-Guided Watermarking for LLMs: A Test-Time Framework for Robust and Traceable Text Generation16 Apr 2025 0 repositories listed
-
Rethinking the Generation of High-Quality CoT Data from the Perspective of LLM-Adaptive Question Difficulty Grading16 Apr 2025 0 repositories listed
-
Heimdall: test-time scaling on the generative verification14 Apr 2025 0 repositories listed
-
GPT Carry-On: Training Foundation Model for Customization Could Be Simple, Scalable and Affordable10 Apr 2025 0 repositories listed
-
Supervised Optimism Correction: Be Confident When LLMs Are Sure10 Apr 2025 0 repositories listed
-
MDIT: A Model-free Data Interpolation Method for Diverse Instruction Tuning9 Apr 2025 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.