Browse State-of-the-Art › GSM8K
GSM8K
209 papers with code · 1 benchmark · 1 dataset archive 2025-07-28
Benchmarks archive 2025-07-28
2 leaderboard tables shown for this task, 1 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| GSM8K (3 rows) | Xolver | Xolver: Multi-Agent Reasoning with Holistic Experience Learning... | code | — | Compare |
| gsm8k (5-shots) (0 rows) | no rows in the archive | — | — | ||
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
1 dataset whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 209 papers with code (439 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 Jan 2022 19 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedWe explore how generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning.
-
18 Jun 2024 7 repositories listed Syntology ran 15 of 29 samples · 14 unverifiedWe introduce ChatGLM, an evolving family of large language models that we have been developing over time.
-
15 Jul 2024 6 repositories listedThis report introduces the Qwen2 series, the latest addition to our large language models and large multimodal models.
-
27 Oct 2021 6 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedState-of-the-art language models can match human performance on many tasks, but they still struggle to robustly perform multi-step mathematical reasoning.
-
7 Sep 2023 4 repositories listed Syntology ran 12 of 14 samples · 2 unverified · 2 pointer-only (licence)In this work, we propose Optimization by PROmpting (OPRO), a simple and effective approach to leverage large language models (LLMs) as optimizers, where the optimization task is described in natural language.
-
6 Oct 2022 4 repositories listedFinally, we show that the multilingual reasoning abilities of language models extend to other tasks such as commonsense reasoning and word-in-context semantic judgment.
-
24 May 2022 4 repositories listed Syntology ran 0 of 4 samples · 4 unverified · 1 pointer-only (licence)Pretrained large language models (LLMs) are widely used in many sub-fields of natural language processing (NLP) and generally known as excellent few-shot learners with task-specific exemplars.
-
19 Sep 2024 3 repositories listedThis method constructs the low-rank update matrix through the replication of these shared, partitioned shards.
-
12 Aug 2024 3 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedThis paper introduces rStar, a self-play mutual reasoning approach that significantly improves reasoning capabilities of small language models (SLMs) without fine-tuning or superior models.
-
14 Dec 2023 3 repositories listedIn this paper, we present an innovative process-oriented math process reward model called \textbf{Math-Shepherd}, which assigns a reward score to each step of math problem solutions.
-
6 Nov 2023 3 repositories listed Syntology ran 2 of 3 samples · 1 unverified · 3 pointer-only (licence)We experiment with encoder- and decoder-based LMs, showing that: (1) SFT delta parameter value ranges are typically small (within 0.
-
29 Aug 2023 3 repositories listedDevelopers face decisions regarding the use of LLMs for directly performing tasks within applications as well as for generating and executing code to accomplish these tasks.
-
27 May 2023 3 repositories listed Syntology ran 5 of 5 samples · 0 unverified · 5 pointer-only (licence)Inspired by this framework, we introduce Matrix-SSL, a novel approach that leverages matrix information theory to interpret the maximum entropy encoding loss as matrix uniformity loss.
-
18 Nov 2022 3 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedMuch of this success can be attributed to prompting methods such as "chain-of-thought'', which employ LLMs for both understanding the problem description by decomposing it into steps, as well as solving each step of the…
-
21 Mar 2022 3 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedChain-of-thought prompting combined with pre-trained large language models has achieved encouraging results on complex reasoning tasks.
-
20 May 2025 2 repositories listed Syntology ran 2 of 2 samples · 0 unverifiedLarge reasoning models (LRMs), such as OpenAI o1 and DeepSeek-R1, have significantly enhanced their reasoning capabilities by generating longer chains of thought, demonstrating outstanding performance across a variety…
-
20 Dec 2024 2 repositories listed Syntology ran 2 of 3 samples · 1 unverifiedWhile Direct Preference Optimization (DPO) has shown promise in aligning LLMs with human preferences, it is less suitable for multi-step reasoning tasks because (1) DPO relies on paired preference data, which is not…
-
11 Oct 2024 2 repositories listed Syntology ran 0 of 1 samples · 1 unverified · 1 pointer-only (licence)Large language models (LLMs) like GPT-4, PaLM, and LLaMA have shown significant improvements in various reasoning tasks.
-
11 Oct 2024 2 repositories listedReinforcement Learning (RL) plays a crucial role in aligning large language models (LLMs) with human preferences and improving their ability to perform complex tasks.
-
10 Oct 2024 2 repositories listed Syntology ran 4 of 19 samples · 15 unverified · 19 pointer-only (licence)However, existing benchmarks like GSM8K or MATH are now being solved with high accuracy (e.
-
7 Oct 2024 2 repositories listedWhile the performance of LLMs on GSM8K has significantly improved in recent years, it remains unclear whether their mathematical reasoning capabilities have genuinely advanced, raising questions about the reliability of…
-
31 Jul 2024 2 repositories listed Syntology ran 4 of 5 samples · 1 unverified · 5 pointer-only (licence)Across multiple tasks and models, we observe that coverage -- the fraction of problems that are solved by any generated sample -- scales with the number of samples over four orders of magnitude.
-
1 May 2024 2 repositories listed Syntology ran 1 of 2 samples · 1 unverifiedWe introduce an approach aimed at enhancing the reasoning capabilities of Large Language Models (LLMs) through an iterative preference learning process inspired by the successful strategy employed by AlphaZero.
-
14 Mar 2024 2 repositories listed Syntology ran 4 of 9 samples · 5 unverifiedCrucially, these improvements require no fine-tuning on these tasks.
-
7 Mar 2024 2 repositories listed Syntology ran 6 of 28 samples · 22 unverified · 12 pointer-only (licence)This paper shows that the LLaMA-2 7B model with common pre-training already exhibits strong mathematical abilities, as evidenced by its impressive accuracy of 97.
-
12 Feb 2024 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We release our curated AutoMathText dataset to facilitate future research in automated domain-specific data curation.
-
16 Jan 2024 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedLarge language models (LLMs) have seen considerable advancements in natural language understanding tasks, yet there remains a gap to bridge before attaining true artificial general intelligence, especially concerning…
-
28 Dec 2023 2 repositories listedIn this work, we introduce a novel evaluation paradigm for Large Language Models (LLMs) that compels them to transition from a traditional question-answering role, akin to a student, to a solution-scoring role, akin to…
-
10 Nov 2023 2 repositories listed Syntology ran 3 of 3 samples · 0 unverifiedWe propose the Data Contamination Quiz (DCQ), a simple and effective approach to detect data contamination in large language models (LLMs) and estimate the amount of it.
-
31 Oct 2023 2 repositories listed Syntology ran 7 of 11 samples · 4 unverifiedThis indicates that crafting multilingual corpora can be regarded as a vital strategy for enhancing model performance in a specific language, especially in mathematical reasoning tasks.
Syntology lines on 22 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections