Browse State-of-the-Art › Math › Papers, page 11
Math
Papers archive 2025-07-28
archive papers tagged: 1,596 · with a code link: 765 · where Syntology ran a sample: 349 (286 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (349 of 1,596 tagged: 286 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument)
Page 11 of 16: papers 1,001 to 1,100 of 1,596, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning25 Jan 2025 0 repositories listed
-
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images24 Jan 2025 0 repositories listed
-
Advancing Mathematical Reasoning in Language Models: The Impact of Problem-Solving Data, Data Synthesis Methods, and Training Stages23 Jan 2025 0 repositories listed
-
An Optimal Transport approach to arbitrage correction: Application to volatility Stress-Tests21 Jan 2025 0 repositories listed
-
Is your LLM trapped in a Mental Set? Investigative study on how mental sets affect the reasoning capabilities of LLMs21 Jan 2025 0 repositories listed
-
RedStar: Does Scaling Long-CoT Data Unlock Better Slow-Reasoning Systems?20 Jan 2025 0 repositories listed
-
Chain-of-Reasoning: Towards Unified Mathematical Reasoning in Large Language Models via a Multi-Paradigm Perspective19 Jan 2025 0 repositories listed
-
Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback18 Jan 2025 0 repositories listed
-
Cascaded Self-Evaluation Augmented Training for Efficient Multimodal Large Language Models10 Jan 2025 0 repositories listed
-
A General Retrieval-Augmented Generation Framework for Multimodal Case-Based Reasoning Applications9 Jan 2025 0 repositories listed
-
End-to-End Bangla AI for Solving Math Olympiad Problem Benchmark: Leveraging Large Language Model Using Integrated Approach8 Jan 2025 0 repositories listed
-
InfiFusion: A Unified Framework for Enhanced Cross-Model Reasoning via LLM Fusion6 Jan 2025 0 repositories listed
-
Quantization Meets Reasoning: Exploring LLM Low-Bit Quantization Degradation for Mathematical Reasoning6 Jan 2025 0 repositories listed
-
Empowering Bengali Education with AI: Solving Bengali Math Word Problems through Transformer Models5 Jan 2025 0 repositories listed
-
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap5 Jan 2025 0 repositories listed
-
Instruction-Following Pruning for Large Language Models3 Jan 2025 0 repositories listed
-
Recursive Decomposition of Logical Thoughts: Framework for Superior Reasoning and Knowledge Propagation in Large Language Models3 Jan 2025 0 repositories listed
-
Experimental Demonstration of an Optical Neural PDE Solver via On-Chip PINN Training1 Jan 2025 0 repositories listed
-
Rethink Delay Doppler Channels and Time-Frequency Coding31 Dec 2024 0 repositories listed
-
Measuring Large Language Models Capacity to Annotate Journalistic Sourcing30 Dec 2024 0 repositories listed
-
Slow Perception: Let's Perceive Geometric Figures Step-by-step30 Dec 2024 0 repositories listed
-
Dynamic Skill Adaptation for Large Language Models26 Dec 2024 0 repositories listed
-
Evaluating the Design Features of an Intelligent Tutoring System for Advanced Mathematics Learning23 Dec 2024 0 repositories listed
-
StructTest: Benchmarking LLMs' Reasoning through Compositional Structured Outputs23 Dec 2024 0 repositories listed
-
Towards Intrinsic Self-Correction Enhancement in Monte Carlo Tree Search Boosted Reasoning via Iterative Preference Learning23 Dec 2024 0 repositories listed
-
Ask-Before-Detection: Identifying and Mitigating Conformity Bias in LLM-Powered Error Detector for Math Word Problem Solutions22 Dec 2024 0 repositories listed
-
System-2 Mathematical Reasoning via Enriched Instruction Tuning22 Dec 2024 0 repositories listed
-
Correct implied volatility shapes and reliable pricing in the rough Heston model20 Dec 2024 0 repositories listed
-
Ensembling Large Language Models with Process Reward-Guided Tree Search for Better Complex Reasoning20 Dec 2024 0 repositories listed
-
Formal Mathematical Reasoning: A New Frontier in AI20 Dec 2024 0 repositories listed
-
AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling19 Dec 2024 0 repositories listed
-
Conceptual In-Context Learning and Chain of Concepts: Solving Complex Conceptual Problems Using Large Language Models19 Dec 2024 0 repositories listed
-
Data for Mathematical Copilots: Better Ways of Presenting Proofs for Machine Learning19 Dec 2024 0 repositories listed
-
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models18 Dec 2024 0 repositories listed
-
Strictly monotone mean-variance preferences with applications to portfolio selection18 Dec 2024 0 repositories listed
-
LinguaLIFT: An Effective Two-stage Instruction Tuning Framework for Low-Resource Language Tasks17 Dec 2024 0 repositories listed
-
A Survey of Mathematical Reasoning in the Era of Multimodal Large Language Model: Benchmark, Method & Challenges16 Dec 2024 0 repositories listed
-
A Graph-Based Synthetic Data Pipeline for Scaling High-Quality Reasoning Instructions12 Dec 2024 0 repositories listed
-
Dipper: Diversity in Prompts for Producing Large Language Model Ensembles in Reasoning tasks12 Dec 2024 0 repositories listed
-
Geo-LLaVA: A Large Multi-Modal Model for Solving Geometry Math Problems with Meta In-Context Learning12 Dec 2024 0 repositories listed
-
Learning to Solve Domain-Specific Calculation Problems with Knowledge-Intensive Programs Generator12 Dec 2024 0 repositories listed
-
Greek2MathTex: A Greek Speech-to-Text Framework for LaTeX Equations Generation11 Dec 2024 0 repositories listed
-
MNIST-Fraction: Enhancing Math Education with AI-Driven Fraction Detection and Analysis11 Dec 2024 0 repositories listed
-
Mining Math Conjectures from LLMs: A Pruning Approach9 Dec 2024 0 repositories listed
-
When Dimensionality Reduction Meets Graph (Drawing) Theory: Introducing a Common Framework, Challenges and Opportunities9 Dec 2024 0 repositories listed
-
Chimera: Improving Generalist Model with Domain-Specific Experts8 Dec 2024 0 repositories listed
-
Hard Math -- Easy UVM: Pragmatic solutions for verifying hardware algorithms using UVM6 Dec 2024 0 repositories listed
-
Neuro-Symbolic Data Generation for Math Reasoning6 Dec 2024 0 repositories listed
-
Automated LaTeX Code Generation from Handwritten Math Expressions Using Vision Transformer5 Dec 2024 0 repositories listed
-
Enhancing Mathematical Reasoning in LLMs with Background Operators5 Dec 2024 0 repositories listed
-
RedStone: Curating General, Code, Math, and QA Data for Large Language Models4 Dec 2024 0 repositories listed
-
MALT: Improving Reasoning with Multi-Agent LLM Training2 Dec 2024 0 repositories listed
-
Yi-Lightning Technical Report2 Dec 2024 0 repositories listed
-
Reverse Thinking Makes LLMs Stronger Reasoners29 Nov 2024 0 repositories listed
-
A Lean Dataset for International Math Olympiad: Small Steps towards Writing Math Proofs for Hard Problems28 Nov 2024 0 repositories listed
-
Mars-PO: Multi-Agent Reasoning System Preference Optimization28 Nov 2024 0 repositories listed
-
Embracing AI in Education: Understanding the Surge in Large Language Model Use by Secondary Students27 Nov 2024 0 repositories listed
-
Learning by Analogy: Enhancing Few-Shot Prompting for Math Word Problem Solving with Computational Graph-Based Retrieval25 Nov 2024 0 repositories listed
-
Unraveling Arithmetic in Large Language Models: The Role of Algebraic Structures25 Nov 2024 0 repositories listed
-
Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training21 Nov 2024 0 repositories listed
-
OpenAI-o1 AB Testing: Does the o1 model really do good reasoning in math problem solving?9 Nov 2024 0 repositories listed
-
VISTA: Visual Integrated System for Tailored Automation in Math Problem Generation Using LLM8 Nov 2024 0 repositories listed
-
Evaluating GPT-4 at Grading Handwritten Solutions in Math Exams7 Nov 2024 0 repositories listed
-
Self-Consistency Preference Optimization6 Nov 2024 0 repositories listed
-
Automatic Generation of Question Hints for Mathematics Problems using Large Language Models in Educational Technology5 Nov 2024 0 repositories listed
-
Dictionary Insertion Prompting for Multilingual Reasoning on Multilingual Large Language Models2 Nov 2024 0 repositories listed
-
Automated Feedback in Math Education: A Comparative Analysis of LLMs for Open-Ended Responses29 Oct 2024 0 repositories listed
-
DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models29 Oct 2024 0 repositories listed
-
Improving Math Problem Solving in Large Language Models Through Categorization and Strategy Tailoring29 Oct 2024 0 repositories listed
-
28 Oct 2024 0 repositories listed Syntology 0 ran · 1 unverified (of 1 harvested sample) · 1 pointer-only (licence)
-
Can Stories Help LLMs Reason? Curating Information Space Through Narrative25 Oct 2024 0 repositories listed
-
From Blind Solvers to Logical Thinkers: Benchmarking LLMs' Logical Integrity on Faulty Mathematical Problems24 Oct 2024 0 repositories listed
-
Mixture of Parrots: Experts improve memorization more than reasoning24 Oct 2024 0 repositories listed
-
ReasonAgain: Using Extractable Symbolic Programs to Evaluate Mathematical Reasoning24 Oct 2024 0 repositories listed
-
MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning23 Oct 2024 0 repositories listed
-
Forewarned is Forearmed: Leveraging LLMs for Data Synthesis through Failure-Inducing Exploration22 Oct 2024 0 repositories listed
-
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation22 Oct 2024 0 repositories listed
-
Optimizing Chain-of-Thought Reasoning: Tackling Arranging Bottleneck via Plan Augmentation22 Oct 2024 0 repositories listed
-
Polyak's Heavy Ball Method Achieves Accelerated Local Rate of Convergence under Polyak-Lojasiewicz Inequality22 Oct 2024 0 repositories listed
-
No more hard prompts: SoftSRV prompting for synthetic data generation21 Oct 2024 0 repositories listed
-
PromptHive: Bringing Subject Matter Experts Back to the Forefront with Collaborative Prompt Engineering for Educational Content Creation21 Oct 2024 0 repositories listed
-
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology19 Oct 2024 0 repositories listed
-
On Designing Effective RL Reward at Training Time for LLM Reasoning19 Oct 2024 0 repositories listed
-
Bridging the Training-Inference Gap in LLMs by Leveraging Self-Generated Tokens18 Oct 2024 0 repositories listed
-
LLM The Genius Paradox: A Linguistic and Math Expert's Struggle with Simple Word-based Counting Problems18 Oct 2024 0 repositories listed
-
Step Guided Reasoning: Improving Mathematical Reasoning using Guidance Generation and Step Reasoning18 Oct 2024 0 repositories listed
-
When Not to Answer: Evaluating Prompts on GPT Models for Effective Abstention in Unanswerable Math Word Problems16 Oct 2024 0 repositories listed
-
MIND: Math Informed syNthetic Dialogues for Pretraining LLMs15 Oct 2024 0 repositories listed
-
Speculative Knowledge Distillation: Bridging the Teacher-Student Gap Through Interleaved Sampling15 Oct 2024 0 repositories listed
-
Embedding Self-Correction as an Inherent Ability in Large Language Models for Enhanced Mathematical Reasoning14 Oct 2024 0 repositories listed
-
Innovative Thinking, Infinite Humor: Humor Research of Large Language Models through Structured Thought Leaps14 Oct 2024 0 repositories listed
-
Dualformer: Controllable Fast and Slow Thinking by Learning with Randomized Reasoning Traces13 Oct 2024 0 repositories listed
-
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning13 Oct 2024 0 repositories listed
-
Testing GPT-4-o1-preview on math and science problems: A follow-up study11 Oct 2024 0 repositories listed
-
Cognitive Noise and Altruistic Preferences10 Oct 2024 0 repositories listed
-
Hallucinating AI Hijacking Attack: Large Language Models and Malicious Code Recommenders9 Oct 2024 0 repositories listed
-
Herald: A Natural Language Annotated Lean 4 Dataset9 Oct 2024 0 repositories listed
-
Subtle Errors Matter: Preference Learning via Error-injected Self-editing9 Oct 2024 0 repositories listed
-
Beyond Captioning: Task-Specific Prompting for Improved VLM Performance in Mathematical Reasoning8 Oct 2024 0 repositories listed
Syntology lines on 2 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.