Browse State-of-the-Art › Math › Papers, page 12
Math
Papers archive 2025-07-28
archive papers tagged: 1,596 · with a code link: 765 · where Syntology ran a sample: 349 (286 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument) Syntology
Show: all tagged papersonly where code ran (349 of 1,596 tagged: 286 with a run with no instrument failure, 63 where every run was a failure of Syntology's instrument)
Page 12 of 16: papers 1,101 to 1,200 of 1,596, in archive order: by repositories listed in the archive (most first), then newest first, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers that list no repository come after every paper that lists one.
Papers without a page here are shown as plain text. A Syntology line reads “N ran (of which C constructed an object rather than computing a result; K with no instrument failure: H honoured, V violated, P with no contract checked; I where Syntology's instrument failed) · U unverified”; the figure “where Syntology's instrument failed” counts failures of Syntology's instrument, not of the code. When the archive marks a repository official for the paper, the line starts with that repository's state (the archive's flag, not a verdict on who wrote the code); hover it for the repositories the samples that ran came from. Abstracts are on each paper's page.
-
FG-PRM: Fine-grained Hallucination Detection and Mitigation in Language Model Mathematical Reasoning8 Oct 2024 0 repositories listed
-
Solving Functional Optimization with Deep Networks and Variational Principles8 Oct 2024 0 repositories listed
-
fPLSA: Learning Semantic Structures in Document Collections Using Foundation Models7 Oct 2024 0 repositories listed
-
Intriguing Properties of Large Language and Vision Models7 Oct 2024 0 repositories listed
-
Reasoning Paths Optimization: Learning to Reason and Explore From Diverse Paths7 Oct 2024 0 repositories listed
-
Rule-based Data Selection for Large Language Models7 Oct 2024 0 repositories listed
-
BloomWise: Enhancing Problem-Solving capabilities of Large Language Models using Bloom's-Taxonomy-Inspired Prompts5 Oct 2024 0 repositories listed
-
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification5 Oct 2024 0 repositories listed
-
Deliberate Reasoning for LLMs as Structure-aware Planning with Accurate World Model4 Oct 2024 0 repositories listed
-
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation3 Oct 2024 0 repositories listed
-
CodePMP: Scalable Preference Model Pretraining for Large Language Model Reasoning3 Oct 2024 0 repositories listed
-
Geometry is All You Need: A Unified Taxonomy of Matrix and Tensor Factorization for Compression of Generative Language Models3 Oct 2024 0 repositories listed
-
Can We Further Elicit Reasoning in LLMs? Critic-Guided Planning with Retrieval-Augmentation for Solving Challenging Tasks2 Oct 2024 0 repositories listed
-
Deep Knowledge Tracing for Personalized Adaptive Learning at Historically Black Colleges and Universities2 Oct 2024 0 repositories listed
-
Evaluating Robustness of Reward Models for Mathematical Reasoning2 Oct 2024 0 repositories listed
-
Layer Swapping for Zero-Shot Cross-Lingual Transfer in Large Language Models2 Oct 2024 0 repositories listed
-
Not All LLM Reasoners Are Created Equal2 Oct 2024 0 repositories listed
-
PersonaMath: Enhancing Math Reasoning through Persona-Driven Data Augmentation2 Oct 2024 0 repositories listed
-
Step-by-Step Reasoning for Math Problems via Twisted Sequential Monte Carlo2 Oct 2024 0 repositories listed
-
Instance-adaptive Zero-shot Chain-of-Thought Prompting30 Sep 2024 0 repositories listed
-
The Perfect Blend: Redefining RLHF with Mixture of Judges30 Sep 2024 0 repositories listed
-
INC-Math: Integrating Natural Language and Code for Enhanced Mathematical Reasoning in Large Language Models28 Sep 2024 0 repositories listed
-
On the Inductive Bias of Stacking Towards Improving Reasoning27 Sep 2024 0 repositories listed
-
Revisiting the Superficial Alignment Hypothesis27 Sep 2024 0 repositories listed
-
Learning to Love Edge Cases in Formative Math Assessment: Using the AMMORE Dataset and Chain-of-Thought Prompting to Improve Grading Accuracy26 Sep 2024 0 repositories listed
-
Democratizing Signal Processing and Machine Learning: Math Learning Equity for Elementary and Middle School Students25 Sep 2024 0 repositories listed
-
LLaMa-SciQ: An Educational Chatbot for Answering Science MCQ25 Sep 2024 0 repositories listed
-
Models Can and Should Embrace the Communicative Nature of Human-Generated Math25 Sep 2024 0 repositories listed
-
PMSS: Pretrained Matrices Skeleton Selection for LLM Fine-tuning25 Sep 2024 0 repositories listed
-
ControlMath: Controllable Data Generation Promotes Math Generalist Models20 Sep 2024 0 repositories listed
-
InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning19 Sep 2024 0 repositories listed
-
GRIN: GRadient-INformed MoE18 Sep 2024 0 repositories listed
-
18 Sep 2024 0 repositories listed
-
NVLM: Open Frontier-Class Multimodal LLMs17 Sep 2024 0 repositories listed
-
GPT takes the SAT: Tracing changes in Test Difficulty and Math Performance of Students16 Sep 2024 0 repositories listed
-
CPL: Critical Plan Step Learning Boosts LLM Generalization in Reasoning Tasks13 Sep 2024 0 repositories listed
-
Cracking the Code: Multi-domain LLM Evaluation on Real-World Professional Exams in Indonesia13 Sep 2024 0 repositories listed
-
Alignment with Preference Optimization Is All You Need for LLM Safety12 Sep 2024 0 repositories listed
-
Knowledge Tagging with Large Language Model based Multi-Agent System12 Sep 2024 0 repositories listed
-
Leveraging Unstructured Text Data for Federated Instruction Tuning of Large Language Models11 Sep 2024 0 repositories listed
-
A Practice of Post-Training on Llama-3 70B with Optimal Selection of Additional Language Mixture Ratio10 Sep 2024 0 repositories listed
-
Building Math Agents with Multi-Turn Iterative Preference Learning4 Sep 2024 0 repositories listed
-
Deconfounded Causality-aware Parameter-Efficient Fine-Tuning for Problem-Solving Improvement of LLMs4 Sep 2024 0 repositories listed
-
Prompt Baking4 Sep 2024 0 repositories listed
-
Wavelet GPT: Wavelet Inspired Large Language Models4 Sep 2024 0 repositories listed
-
S³c-Math: Spontaneous Step-level Self-correction Makes Large Language Models Better Mathematical Reasoners3 Sep 2024 0 repositories listed
-
Critic-CoT: Boosting the reasoning abilities of large language model via Chain-of-thoughts Critic29 Aug 2024 0 repositories listed
-
Entropic Distribution Matching in Supervised Fine-tuning of LLMs: Less Overfitting and Better Diversity29 Aug 2024 0 repositories listed
-
Logic Contrastive Reasoning with Lightweight Large Language Model for Math Word Problems29 Aug 2024 0 repositories listed
-
Physics of Language Models: Part 2.2, How to Learn From Mistakes on Grade-School Math Problems29 Aug 2024 0 repositories listed
-
SIaM: Self-Improving Code-Assisted Mathematical Reasoning of Large Language Models28 Aug 2024 0 repositories listed
-
Generative Verifiers: Reward Modeling as Next-Token Prediction27 Aug 2024 0 repositories listed
-
Students' Perceived Roles, Opportunities, and Challenges of a Generative AI-powered Teachable Agent: A Case of Middle School Math Class26 Aug 2024 0 repositories listed
-
Multi-tool Integration Application for Math Reasoning Using Large Language Model22 Aug 2024 0 repositories listed
-
Mathematical Information Retrieval: Search and Question Answering21 Aug 2024 0 repositories listed
-
QPO: Query-dependent Prompt Optimization via Multi-Loop Offline Reinforcement Learning20 Aug 2024 0 repositories listed
-
A Study of PHOC Spatial Region Configurations for Math Formula Retrieval17 Aug 2024 0 repositories listed
-
Large Language Models Might Not Care What You Are Saying: Prompt Format Beats Descriptions16 Aug 2024 0 repositories listed
-
Does Reasoning Emerge? Examining the Probabilities of Causation in Large Language Models15 Aug 2024 0 repositories listed
-
A Perspective on Large Language Models, Intelligent Machines, and Knowledge Acquisition13 Aug 2024 0 repositories listed
-
P3: A Policy-Driven, Pace-Adaptive, and Diversity-Promoted Framework for data pruning in LLM Training10 Aug 2024 0 repositories listed
-
Examining the Behavior of LLM Architectures Within the Framework of Standardized National Exams in Brazil9 Aug 2024 0 repositories listed
-
AltCanvas: A Tile-Based Image Editor with Generative AI for Blind or Visually Impaired People5 Aug 2024 0 repositories listed
-
The Logic of Political Survival Revisited: Consequences of Elite Uncertainty Under Authoritarian Rule4 Aug 2024 0 repositories listed
-
Towards Effective and Efficient Continual Pre-training of Large Language Models26 Jul 2024 0 repositories listed
-
Recursive Introspection: Teaching Language Model Agents How to Self-Improve25 Jul 2024 0 repositories listed
-
A LLM Benchmark based on the Minecraft Builder Dialog Agent Task17 Jul 2024 0 repositories listed
-
CCoE: A Compact LLM with Collaboration of Experts16 Jul 2024 0 repositories listed
-
Reasoning with Large Language Models, a Survey16 Jul 2024 0 repositories listed
-
TelecomGPT: A Framework to Build Telecom-Specfic Large Language Models12 Jul 2024 0 repositories listed
-
Token-Supervised Value Models for Enhancing Mathematical Reasoning Capabilities of Large Language Models12 Jul 2024 0 repositories listed
-
Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist11 Jul 2024 0 repositories listed
-
Skywork-Math: Data Scaling Laws for Mathematical Reasoning in Large Language Models -- The Story Goes On11 Jul 2024 0 repositories listed
-
ConvNLP: Image-based AI Text Detection9 Jul 2024 0 repositories listed
-
Advancing Process Verification for Large Language Models via Tree-Based Preference Learning29 Jun 2024 0 repositories listed
-
CMMaTH: A Chinese Multi-modal Math Skill Evaluation Benchmark for Foundation Models28 Jun 2024 0 repositories listed
-
ScaleBiO: Scalable Bilevel Optimization for LLM Data Reweighting28 Jun 2024 0 repositories listed
-
Task Oriented In-Domain Data Augmentation24 Jun 2024 0 repositories listed
-
Generative AI for Enhancing Active Learning in Education: A Comparative Study of GPT-3.5 and GPT-4 in Crafting Customized Test Questions20 Jun 2024 0 repositories listed
-
Q*: Improving Multi-step Reasoning for LLMs with Deliberative Planning20 Jun 2024 0 repositories listed
-
Knowledge Tagging System on Math Questions via LLMs with Flexible Demonstration Retriever19 Jun 2024 0 repositories listed
-
Navigating the Labyrinth: Evaluating and Enhancing LLMs' Ability to Reason About Search Problems18 Jun 2024 0 repositories listed
-
Program Synthesis Benchmark for Visual Programming in XLogoOnline Environment17 Jun 2024 0 repositories listed
-
Self-MoE: Towards Compositional Large Language Models with Self-Specialized Experts17 Jun 2024 0 repositories listed
-
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning16 Jun 2024 0 repositories listed
-
CLST: Cold-Start Mitigation in Knowledge Tracing by Aligning a Generative Language Model as a Students' Knowledge Tracer13 Jun 2024 0 repositories listed
-
ReMI: A Dataset for Reasoning with Multiple Images13 Jun 2024 0 repositories listed
-
Can I understand what I create? Self-Knowledge Evaluation of Large Language Models10 Jun 2024 0 repositories listed
-
Human Learning about AI8 Jun 2024 0 repositories listed
-
A multi-core periphery perspective: Ranking via relative centrality6 Jun 2024 0 repositories listed
-
Improve Mathematical Reasoning in Language Models by Automated Process Supervision5 Jun 2024 0 repositories listed
-
D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language Models3 Jun 2024 0 repositories listed
-
Code Pretraining Improves Entity Tracking Abilities of Language Models31 May 2024 0 repositories listed
-
Divide-and-Conquer Meets Consensus: Unleashing the Power of Functions in Code Generation30 May 2024 0 repositories listed
-
Arithmetic Reasoning with LLM: Prolog Generation & Permutation28 May 2024 0 repositories listed
-
MindStar: Enhancing Math Reasoning in Pre-trained LLMs at Inference Time25 May 2024 0 repositories listed
-
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs24 May 2024 0 repositories listed
-
Large Language Models Can Self-Correct with Key Condition Verification23 May 2024 0 repositories listed
-
"Turing Tests" For An AI Scientist22 May 2024 0 repositories listed
Syntology lines on 1 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced; each line links to that paper's sample list. Syntology's record for this page has not changed since , the first build that kept a record date for it; when this build read Syntology's graph is in the build record. For agents: get_harvested_code_for_paper(arxiv_id) lists each paper's samples; how to connect.