Browse State-of-the-Art › Mathematical Problem-Solving
Mathematical Problem-Solving
55 papers with code · 0 benchmarks · 2 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
2 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 55 papers with code (106 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
5 Mar 2021 5 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics.
-
4 Nov 2024 3 repositories listedIn this paper, we introduce Hunyuan-Large, which is currently the largest open-source Transformer-based mixture of experts model, with a total of 389 billion parameters and 52 billion activation parameters, capable of…
-
11 Jul 2024 3 repositories listedThe mathematical capabilities of Multi-modal Large Language Models (MLLMs) remain under-explored with three areas to be improved: visual encoding of math diagrams, diagram-language alignment, and chain-of-thought (CoT)…
-
26 Jun 2024 3 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedThis paper investigates the mathematical problem-solving capabilities of LLMs using the newly developed "MathOdyssey" dataset.
-
3 Apr 2024 3 repositories listed Syntology ran 1 of 4 samples · 3 unverifiedLarge language models (LLMs) have shown excellent mastering of human language, but still struggle in real-world applications that require mathematical problem-solving.
-
18 Dec 2023 3 repositories listed Syntology ran 4 of 6 samples · 2 unverified · 6 pointer-only (licence)We first analyze the limitations of current Multimodal Large Language Models (MLLMs) in this area: they struggle to accurately comprehending basic geometric elements and their relationships.
-
9 Jun 2025 2 repositories listed Syntology ran 9 of 10 samples · 1 unverifiedInequality proving, crucial across diverse scientific and mathematical fields, tests advanced reasoning skills such as discovering tight bounds and strategic theorem application.
-
12 May 2025 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedLarge Language Models (LLMs) often struggle with mathematical reasoning tasks requiring precise, verifiable computation.
-
4 Jul 2025 1 repository listed Syntology ran 0 of 20 samples · 20 unverified · 16 pointer-only (licence)Multi-agent systems (MAS) have emerged as a powerful paradigm for orchestrating large language models (LLMs) and specialized tools to collaboratively address complex tasks.
-
16 Jun 2025 1 repository listed Syntology ran 0 of 11 samples · 11 unverified · 11 pointer-only (licence)Recent advances in large language models (LLMs), particularly those enhanced through reinforced post-training, have demonstrated impressive reasoning capabilities, as exemplified by models such as OpenAI o1 and…
-
10 Jun 2025 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)A prerequisite for the scalability of RLVR is a high-quality problem set with precise and verifiable answers.
-
8 Jun 2025 1 repository listedLarge Language Models (LLMs) have achieved remarkable success in tasks requiring complex reasoning, such as code generation, mathematical problem solving, and algorithmic synthesis -- especially when aided by reasoning…
-
5 Jun 2025 1 repository listedDespite rapid advances in vision-language models (VLMs), current benchmarks for multimodal reasoning fall short in three key dimensions.
-
28 May 2025 1 repository listedMathematical reasoning tasks have become prominent benchmarks for assessing the reasoning capabilities of LLMs, especially with reinforcement learning (RL) methods such as GRPO showing significant performance gains.
-
26 May 2025 1 repository listedThe resulting GRPO approach, leveraging format-length surrogate signals, not only matches but surpasses the performance of the standard GRPO algorithm relying on ground truth answers in certain scenarios, achieving 40.
-
26 May 2025 1 repository listedLarge Language Models (LLMs) are prone to hallucination, especially during multi-hop and reasoning-intensive tasks such as mathematical problem solving.
-
23 May 2025 1 repository listedIn addition, RaDeR presents the first dense retriever that outperforms BM25 when queries are Chain-of-Thought reasoning steps, underscoring the critical role of reasoning-based retrieval to augment reasoning language…
-
17 May 2025 1 repository listedLarge language models (LLMs) have shown remarkable progress in mathematical problem-solving, but evaluation has largely focused on problems that have exact analytical solutions or involve formal proofs, often…
-
14 May 2025 1 repository listedParameter-efficient fine-tuning (PEFT) methods have shown promise in adapting large language models, yet existing approaches exhibit counter-intuitive phenomena: integrating router into prompt tuning (PT) increases…
-
2 Apr 2025 1 repository listedThis study investigates the reasoning robustness of large language models (LLMs) on mathematical problem-solving tasks under systematically introduced input perturbations.
-
31 Mar 2025 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedThe mathematical problem-solving capabilities of large language models have become a focal point of research, with growing interests in leveraging self-generated reasoning paths as a promising way to refine and enhance…
-
22 Mar 2025 1 repository listedFuture research should focus on interpretability, integration with domain-specific solvers, and improving the robustness of AI-driven decision-making.
-
21 Mar 2025 1 repository listedReasoning capabilities have significantly improved the performance of vision-language models (VLMs) in domains such as mathematical problem-solving, coding, and visual question-answering.
-
20 Mar 2025 1 repository listed Syntology ran 10 of 18 samples · 8 unverifiedLarge Language Models (LLMs) have shown impressive progress in mathematical reasoning.
-
19 Mar 2025 1 repository listedDespite impressive performance across diverse tasks, Multimodal Large Language Models (MLLMs) have yet to fully demonstrate their potential in visual mathematical problem-solving, particularly in accurately perceiving…
-
26 Feb 2025 1 repository listedIn coding tasks, Nexus-driven MASs achieve a 99% pass rate on HumanEval and a flawless 100% on VerilogEval-Human, outperforming cutting-edge reasoning language models such as o3-mini and DeepSeek-R1.
-
21 Feb 2025 1 repository listed Syntology ran 2 of 2 samples · 0 unverified · 2 pointer-only (licence)Despite strong performance on vision-language tasks, Multimodal Large Language Models (MLLMs) struggle with mathematical problem-solving, with both open-source and state-of-the-art models falling short of human…
-
17 Feb 2025 1 repository listedThis paper introduces Code-Vision, a benchmark designed to evaluate the logical understanding and code generation capabilities of Multimodal Large Language Models (MLLMs).
-
16 Feb 2025 1 repository listedLarge Language Models (LLMs) have demonstrated impressive capabilities in natural language processing tasks, such as text generation and semantic understanding.
-
7 Feb 2025 1 repository listedLarge Language Models (LLMs) have demonstrated impressive reasoning capabilities, yet their performance is highly dependent on the prompting strategy and model scale.
Syntology lines on 12 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections