Browse State-of-the-Art › Math › Papers, page 10
Math
Papers archive 2025-07-28
archive papers tagged: 1,596 · with a code link: 765 · where Syntology ran a sample: 349 (286 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (349 of 1,596 tagged: 286 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument)
Page 10 of 16: papers 901 to 1,000 of 1,596, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Reasoning Models Know When They're Right: Probing Hidden States for Self-Verification7 Apr 2025 0 repositories listed
-
Synthetic Data Generation & Multi-Step RL for Reasoning & Tool Use7 Apr 2025 0 repositories listed
-
Retro-Search: Exploring Untaken Paths for Deeper and Efficient Reasoning6 Apr 2025 0 repositories listed
-
oneDAL Optimization for ARM Scalable Vector Extension: Maximizing Efficiency for High-Performance Data Science5 Apr 2025 0 repositories listed
-
Explain with Visual Keypoints Like a Real Mentor! A Benchmark for Multimodal Solution Explanation4 Apr 2025 0 repositories listed
-
Online Difficulty Filtering for Reasoning Oriented Reinforcement Learning4 Apr 2025 0 repositories listed
-
Cross-Lingual Consistency: A Novel Inference Framework for Advancing Reasoning in Large Language Models2 Apr 2025 0 repositories listed
-
Brains vs. Bytes: Evaluating LLM Proficiency in Olympiad Mathematics1 Apr 2025 0 repositories listed
-
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning1 Apr 2025 0 repositories listed
-
Hawkeye:Efficient Reasoning with Model Collaboration1 Apr 2025 0 repositories listed
-
How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study1 Apr 2025 0 repositories listed
-
Investigating Large Language Models in Diagnosing Students' Cognitive Skills in Math Problem-solving1 Apr 2025 0 repositories listed
-
DebFlow: Automating Agent Creation via Agent Debate31 Mar 2025 0 repositories listed
-
Proof or Bluff? Evaluating LLMs on 2025 USA Math Olympiad27 Mar 2025 0 repositories listed
-
1.4 Million Open-Source Distilled Reasoning Dataset to Empower Large Language Model Training25 Mar 2025 0 repositories listed
-
Gemma 3 Technical Report25 Mar 2025 0 repositories listed
-
Scaling Evaluation-time Compute with Reasoning Models as Process Evaluators25 Mar 2025 0 repositories listed
-
Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking25 Mar 2025 0 repositories listed
-
Activation Functions Considered Harmful: Recovering Neural Network Weights through Controlled Channels24 Mar 2025 0 repositories listed
-
Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling24 Mar 2025 0 repositories listed
-
Teaching LLMs for Step-Level Automatic Math Correction via Reinforcement Learning24 Mar 2025 0 repositories listed
-
Long Is More Important Than Difficult for Training Reasoning Models23 Mar 2025 0 repositories listed
-
MathAgent: Leveraging a Mixture-of-Math-Agent Framework for Real-World Multimodal Mathematical Error Detection23 Mar 2025 0 repositories listed
-
Exploring the Hidden Reasoning Process of Large Language Models by Misleading Them20 Mar 2025 0 repositories listed
-
Tapered Off-Policy REINFORCE: Stable and efficient reinforcement learning for LLMs18 Mar 2025 0 repositories listed
-
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach17 Mar 2025 0 repositories listed
-
Pensez: Less Data, Better Reasoning -- Rethinking French LLM17 Mar 2025 0 repositories listed
-
SPIN-Bench: How Well Do LLMs Plan Strategically and Reason Socially?16 Mar 2025 0 repositories listed
-
Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and Natural Language Data13 Mar 2025 0 repositories listed
-
Conformal Prediction Sets for Deep Generative Models via Reduction to Conformal Regression13 Mar 2025 0 repositories listed
-
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory13 Mar 2025 0 repositories listed
-
Understanding the Logical Capabilities of Large Language Models via Out-of-Context Representation Learning13 Mar 2025 0 repositories listed
-
From Text to Visuals: Using LLMs to Generate Math Diagrams with Vector Graphics10 Mar 2025 0 repositories listed
-
Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning10 Mar 2025 0 repositories listed
-
Decoding the Black Box: Integrating Moral Imagination with Technical AI Governance9 Mar 2025 0 repositories listed
-
InftyThink: Breaking the Length Limits of Long-Context Reasoning in Large Language Models9 Mar 2025 0 repositories listed
-
Symbolic Mixture-of-Experts: Adaptive Skill-based Routing for Heterogeneous Reasoning7 Mar 2025 0 repositories listed
-
Benchmarking Reasoning Robustness in Large Language Models6 Mar 2025 0 repositories listed
-
Better Process Supervision with Bi-directional Rewarding Signals6 Mar 2025 0 repositories listed
-
Compositional Causal Reasoning Evaluation in Language Models6 Mar 2025 0 repositories listed
-
HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks6 Mar 2025 0 repositories listed
-
SOLAR: Scalable Optimization of Large-scale Architecture for Reasoning6 Mar 2025 0 repositories listed
-
START: Self-taught Reasoner with Tools6 Mar 2025 0 repositories listed
-
FANS -- Formal Answer Selection for Natural Language Math Reasoning Using Lean45 Mar 2025 0 repositories listed
-
LEWIS (LayEr WIse Sparsity) -- A Training Free Guided Model Merging Approach5 Mar 2025 0 repositories listed
-
Performance Comparison of Large Language Models on Advanced Calculus Problems5 Mar 2025 0 repositories listed
-
Self-Evolved Preference Optimization for Enhancing Mathematical Reasoning in Small Language Models4 Mar 2025 0 repositories listed
-
Cats Confuse Reasoning LLM: Query Agnostic Adversarial Triggers for Reasoning Models3 Mar 2025 0 repositories listed
-
What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret3 Mar 2025 0 repositories listed
-
MV-MATH: Evaluating Multimodal Math Reasoning in Multi-Visual Contexts28 Feb 2025 0 repositories listed
-
Med-RLVR: Emerging Medical Reasoning from a 3B base model via reinforcement Learning27 Feb 2025 0 repositories listed
-
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution25 Feb 2025 0 repositories listed
-
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning25 Feb 2025 0 repositories listed
-
Learning Decentralized Swarms Using Rotation Equivariant Graph Neural Networks24 Feb 2025 0 repositories listed
-
Reasoning with Latent Thoughts: On the Power of Looped Transformers24 Feb 2025 0 repositories listed
-
DISC: DISC: Dynamic Decomposition Improves LLM Inference Scaling23 Feb 2025 0 repositories listed
-
SBSC: Step-By-Step Coding for Improving Mathematical Olympiad Performance23 Feb 2025 0 repositories listed
-
Inference Computation Scaling for Feature Augmentation in Recommendation Systems22 Feb 2025 0 repositories listed
-
Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning21 Feb 2025 0 repositories listed
-
A Survey on Feedback-based Multi-step Reasoning for Large Language Models on Mathematics20 Feb 2025 0 repositories listed
-
BeamLoRA: Beam-Constraint Low-Rank Adaptation19 Feb 2025 0 repositories listed
-
DiffSampling: Enhancing Diversity and Accuracy in Neural Text Generation19 Feb 2025 0 repositories listed
-
The Self-Improvement Paradox: Can Language Models Bootstrap Reasoning Capabilities without External Scaffolding?19 Feb 2025 0 repositories listed
-
Lean-ing on Quality: How High-Quality Data Beats Diverse Multilingual Data in AutoFormalization18 Feb 2025 0 repositories listed
-
Multi-Step Alignment as Markov Games: An Optimistic Online Gradient Descent Approach with Convergence Guarantees18 Feb 2025 0 repositories listed
-
NaturalReasoning: Reasoning in the Wild with 2.8M Challenging Questions18 Feb 2025 0 repositories listed
-
None of the Others: a General Technique to Distinguish Reasoning from Memorization in Multiple-Choice LLM Evaluation Benchmarks18 Feb 2025 0 repositories listed
-
Thinking Outside the (Gray) Box: A Context-Based Score for Assessing Value and Originality in Neural Text Generation18 Feb 2025 0 repositories listed
-
A Study on Leveraging Search and Self-Feedback for Agent Reasoning17 Feb 2025 0 repositories listed
-
Energy-Conscious LLM Decoding: Impact of Text Generation Strategies on GPU Energy Consumption17 Feb 2025 0 repositories listed
-
Hypothesis-Driven Theory-of-Mind Reasoning for Large Language Models17 Feb 2025 0 repositories listed
-
MathFimer: Enhancing Mathematical Reasoning by Expanding Reasoning Steps through Fill-in-the-Middle Task17 Feb 2025 0 repositories listed
-
Scaling Test-Time Compute Without Verification or RL is Suboptimal17 Feb 2025 0 repositories listed
-
Teaching LLMs According to Their Aptitude: Adaptive Reasoning for Mathematical Problem Solving17 Feb 2025 0 repositories listed
-
Why Vision Language Models Struggle with Visual Arithmetic? Towards Enhanced Chart and Geometry Understanding17 Feb 2025 0 repositories listed
-
Graders should cheat: privileged information enables expert-level automated evaluations16 Feb 2025 0 repositories listed
-
1bit-Merging: Dynamic Quantized Merging for Large Language Models15 Feb 2025 0 repositories listed
-
CRANE: Reasoning with constrained LLM generation13 Feb 2025 0 repositories listed
-
MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency13 Feb 2025 0 repositories listed
-
Interactive Sketchpad: A Multimodal Tutoring System for Collaborative, Visual Problem-Solving12 Feb 2025 0 repositories listed
-
O1 Embedder: Let Retrievers Think Before Action11 Feb 2025 0 repositories listed
-
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations10 Feb 2025 0 repositories listed
-
Evolving LLMs' Self-Refinement Capability via Iterative Preference Optimization8 Feb 2025 0 repositories listed
-
BOLT: Bootstrap Long Chain-of-Thought in Language Models without Distillation6 Feb 2025 0 repositories listed
-
Entropy Adaptive Decoding: Dynamic Model Switching for Efficient Inference5 Feb 2025 0 repositories listed
-
Gold-medalist Performance in Solving Olympiad Geometry with AlphaGeometry25 Feb 2025 0 repositories listed
-
Reasoning-as-Logic-Units: Scaling Test-Time Reasoning in Large Language Models Through Logic Unit Alignment5 Feb 2025 0 repositories listed
-
Premise-Augmented Reasoning Chains Improve Error Identification in Math reasoning with LLMs4 Feb 2025 0 repositories listed
-
SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model4 Feb 2025 0 repositories listed
-
Blink of an eye: a simple theory for feature localization in generative models2 Feb 2025 0 repositories listed
-
Learning Autonomous Code Integration for Math Language Models2 Feb 2025 0 repositories listed
-
Rethinking Mixture-of-Agents: Is Mixing Different Large Language Models Beneficial?2 Feb 2025 0 repositories listed
-
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning31 Jan 2025 0 repositories listed
-
Fairshare Data Pricing via Data Valuation for Large Language Models31 Jan 2025 0 repositories listed
-
Pheromone-based Learning of Optimal Reasoning Paths31 Jan 2025 0 repositories listed
-
PixelWorld: Towards Perceiving Everything as Pixels31 Jan 2025 0 repositories listed
-
Examining the Robustness of Large Language Models across Language Complexity30 Jan 2025 0 repositories listed
-
Token-Hungry, Yet Precise: DeepSeek R1 Highlights the Need for Multi-Step Reasoning Over Speed in MATH30 Jan 2025 0 repositories listed
-
Token-by-Token Regeneration and Domain Biases: A Benchmark of LLMs on Advanced Mathematical Problem-Solving28 Jan 2025 0 repositories listed
-
Error Classification of Large Language Models on Math Word Problems: A Dynamically Adaptive Framework26 Jan 2025 0 repositories listed