Browse State-of-the-Art › Mathematical Reasoning
Mathematical Reasoning
395 papers with code · 10 benchmarks · 26 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
11 leaderboard tables shown for this task, 10 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 11 until expanded.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| AIME24 (9 rows) | Xolver | Xolver: Multi-Agent Reasoning with Holistic Experience Learning... | code | — | Compare |
| FrontierMath (6 rows) | o3 | — | — | — | Compare |
| Lila (IID) (6 rows) | Codex (Few-Shot, 175B) | Lila: A Unified Benchmark for Mathematical Reasoning | code | — | Compare |
| Lila (OOD) (6 rows) | Codex (Few-Shot, 175B) | Lila: A Unified Benchmark for Mathematical Reasoning | code | — | Compare |
| PGPS9K (6 rows) | GOLD | GOLD: Geometry Problem Solver with Natural Language Description | code | — | Compare |
| AMC23 (3 rows) | QWQ-32B-preview | — | — | — | Compare |
| GeoQA (2 rows) | GOLD | GOLD: Geometry Problem Solver with Natural Language Description | code | — | Compare |
| MATH500 (1 row) | Search-o1 | Search-o1: Agentic Search-Enhanced Large Reasoning Models | code | — | Compare |
| UniGeo (1 row) | GOLD | GOLD: Geometry Problem Solver with Natural Language Description | code | — | Compare |
| UniGeo (PRV) (1 row) | GAPS | GAPS: Geometry-Aware Problem Solver | — | — | Compare |
| GSM8K (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
26 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
7 subtasks in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 395 papers with code (805 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
17 Jun 2021 74 repositories listed Syntology ran 34 of 84 samples · 50 unverified · 28 pointer-only (licence)We propose Low-Rank Adaptation, or LoRA, which freezes the pre-trained model weights and injects trainable rank decomposition matrices into each layer of the Transformer architecture, greatly reducing the number of…
-
2 Apr 2019 7 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedThe structured nature of the mathematics domain, covering arithmetic, algebra, probability and calculus, enables the construction of training and test splits designed to clearly illuminate the capabilities and…
-
19 Dec 2024 6 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedIn addition, for hosted solutions, the proprietary models currently include two mixture-of-experts (MoE) variants: Qwen2.
-
10 Oct 2023 6 repositories listed Syntology ran 9 of 11 samples · 2 unverified · 1 pointer-only (licence)We introduce Mistral 7B v0.
-
27 Oct 2021 6 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedState-of-the-art language models can match human performance on many tasks, but they still struggle to robustly perform multi-step mathematical reasoning.
-
5 Feb 2024 5 repositories listed Syntology ran 8 of 24 samples · 16 unverifiedMathematical reasoning poses a significant challenge for language models due to its complex and structured nature.
-
5 Mar 2021 5 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics.
-
14 May 2025 4 repositories listedIn this work, we present Qwen3, the latest version of the Qwen model family.
-
22 Jan 2025 4 repositories listedWe introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1.
-
6 May 2025 3 repositories listed Syntology ran 3 of 12 samples · 9 unverifiedReinforcement learning with verifiable rewards (RLVR) has shown promise in enhancing the reasoning capabilities of large language models by learning directly from outcome-based rewards.
-
27 Feb 2025 3 repositories listedLLM routing is a crucial paradigm that dynamically selects the most suitable large language models from a pool of candidates to process diverse inputs, ensuring optimal resource utilization while maintaining response…
-
10 Feb 2025 3 repositories listedWe train our ReasonFlux-32B model with only 8 GPUs and introduces three innovations: (i) a structured and generic thought template library, containing around 500 high-level thought templates capable of generalizing to…
-
5 Feb 2025 3 repositories listedWhile conventional wisdom suggests that sophisticated reasoning tasks demand extensive training data (>100, 000 examples), we demonstrate that complex mathematical reasoning abilities can be effectively elicited with…
-
11 Jul 2024 3 repositories listedThe mathematical capabilities of Multi-modal Large Language Models (MLLMs) remain under-explored with three areas to be improved: visual encoding of math diagrams, diagram-language alignment, and chain-of-thought (CoT)…
-
26 Jun 2024 3 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedThis paper investigates the mathematical problem-solving capabilities of LLMs using the newly developed "MathOdyssey" dataset.
-
19 Jan 2024 3 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 4 pointer-only (licence)In this paper, we introduce the notion of knowledge fusion for LLMs, aimed at combining the capabilities of existing LLMs and transferring them into a single LLM.
-
14 Dec 2023 3 repositories listedIn this paper, we present an innovative process-oriented math process reward model called \textbf{Math-Shepherd}, which assigns a reward score to each step of math problem solutions.
-
25 Jul 2023 3 repositories listed Syntology ran 4 of 14 samples · 10 unverifiedWith the above challenges in mind, in this paper, we propose FacTool, a task and domain agnostic framework for detecting factual errors of texts generated by large language models (e.
-
30 Mar 2023 3 repositories listedMotivated by how humans refine their written text, we introduce Self-Refine, an approach for improving initial outputs from LLMs through iterative feedback and refinement.
-
22 Mar 2023 3 repositories listed Syntology ran 0 of 10 samples · 10 unverifiedWe contend that (this early version of) GPT-4 is part of a new cohort of LLMs (along with ChatGPT and Google's PaLM for example) that exhibit more general intelligence than previous AI models.
-
18 Nov 2022 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedMuch of this success can be attributed to prompting methods such as "chain-of-thought'', which employ LLMs for both understanding the problem description by decomposing it into steps, as well as solving each step of the…
-
8 Dec 2021 3 repositories listedLanguage modelling provides a step towards intelligent communication systems by harnessing large repositories of written human knowledge to better predict and understand the world.
-
5 Nov 2019 3 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedWe study compositional generalization, viz., the problem of zero-shot generalization to novel compositions of concepts in a domain.
-
8 Jul 2025 2 repositories listed Syntology ran 3 of 12 samples · 9 unverified · 12 pointer-only (licence)The strong performance of Skywork-R1V3 primarily stems from our elaborate post-training RL framework, which effectively activates and enhances the model's reasoning ability, without the need for additional continue…
-
29 May 2025 2 repositories listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)In this work, we propose DeepTheorem, a comprehensive informal theorem-proving framework exploiting natural language to enhance LLM mathematical reasoning.
-
19 May 2025 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedWhile Multimodal Large Language Models (MLLMs) have achieved impressive progress in vision-language understanding, they still struggle with complex multi-step reasoning, often producing logically inconsistent or…
-
12 May 2025 2 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedLarge Language Models (LLMs) often struggle with mathematical reasoning tasks requiring precise, verifiable computation.
-
27 Mar 2025 2 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedIn recent years, the rapid development of large reasoning models has resulted in the saturation of existing benchmarks for evaluating mathematical reasoning, highlighting the urgent need for more challenging and…
-
26 Feb 2025 2 repositories listed Syntology ran 5 of 16 samples · 11 unverifiedWe study self-rewarding reasoning large language models (LLMs), which can simultaneously generate step-by-step reasoning and evaluate the correctness of their outputs during the inference time-without external feedback.
-
31 Jan 2025 2 repositories listed Syntology ran 4 of 4 samples · 0 unverifiedAfter supervised finetuning the Qwen2.
Syntology lines on 21 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections