Browse State-of-the-Art › Math
Math
765 papers with code · 0 benchmarks · 6 datasets archive 2025-07-28
Benchmarks archive 2025-07-28
No benchmark for this task in the archive.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 765 papers with code (1,596 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
28 Jan 2022 19 repositories listed Syntology ran 2 of 7 samples · 5 unverifiedWe explore how generating a chain of thought -- a series of intermediate reasoning steps -- significantly improves the ability of large language models to perform complex reasoning.
-
1 Jun 2023 12 repositories listed Syntology ran 12 of 18 samples · 6 unverified · 2 pointer-only (licence)We propose Activation-aware Weight Quantization (AWQ), a hardware-friendly approach for LLM low-bit weight-only quantization.
-
15 Mar 2023 11 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 1 pointer-only (licence)We report the development of GPT-4, a large-scale, multimodal model which can accept image and text inputs and produce text outputs.
-
29 Aug 2019 8 repositories listedThe encoder is a convolutional neural network (CNN) that transforms images into a group of feature maps.
-
18 Jun 2024 7 repositories listed Syntology ran 15 of 29 samples · 14 unverifiedWe introduce ChatGLM, an evolving family of large language models that we have been developing over time.
-
5 Apr 2022 7 repositories listed Syntology ran 30 of 37 samples · 7 unverifiedTo further our understanding of the impact of scale on few-shot learning, we trained a 540-billion parameter, densely activated, Transformer language model, which we call Pathways Language Model PaLM.
-
19 Dec 2024 6 repositories listed Syntology ran 1 of 3 samples · 2 unverifiedIn addition, for hosted solutions, the proprietary models currently include two mixture-of-experts (MoE) variants: Qwen2.
-
10 Oct 2023 6 repositories listed Syntology ran 9 of 11 samples · 2 unverified · 1 pointer-only (licence)We introduce Mistral 7B v0.
-
9 Jun 2022 6 repositories listed Syntology ran 0 of 3 samples · 3 unverifiedBIG-bench focuses on tasks that are believed to be beyond the capabilities of current language models.
-
27 Oct 2021 6 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedState-of-the-art language models can match human performance on many tasks, but they still struggle to robustly perform multi-step mathematical reasoning.
-
3 Feb 2025 5 repositories listed Syntology ran 2 of 17 samples · 15 unverifiedWhile dense rewards also offer an appealing choice for the reinforcement learning (RL) of LLMs since their fine-grained rewards have the potential to address some inherent issues of outcome rewards, such as training…
-
5 Feb 2024 5 repositories listed Syntology ran 8 of 24 samples · 16 unverifiedMathematical reasoning poses a significant challenge for language models due to its complex and structured nature.
-
11 Mar 2021 5 repositories listedWe present a Neural Network based Handwritten Text Recognition (HTR) model architecture that can be trained to recognize full pages of handwritten or printed text without image segmentation.
-
5 Mar 2021 5 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 1 pointer-only (licence)To facilitate future research and increase accuracy on MATH, we also contribute a large auxiliary pretraining dataset which helps teach models the fundamentals of mathematics.
-
5 Feb 2018 5 repositories listedThis paper is an attempt to explain all the matrix calculus you need in order to understand the training of deep neural networks.
-
28 Feb 2025 4 repositories listedMathematical formulas are a fundamental and widely used component in various scientific fields, serving as a universal language for expressing complex concepts and relationships.
-
20 Feb 2025 4 repositories listed Syntology ran 10 of 23 samples · 13 unverifiedInspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in large reasoning models.
-
10 Feb 2025 4 repositories listed Syntology ran 0 of 13 samples · 13 unverifiedLastly, we propose a theory as to why RLSP search strategy is more suitable for LLMs inspired by a remarkable result that says CoT provably increases computational power of LLMs, which grows as the number of steps in…
-
22 Apr 2024 4 repositories listed Syntology ran 8 of 13 samples · 5 unverified · 13 pointer-only (licence)Large language models (LLMs) have been explored in a variety of reasoning tasks including solving of mathematical problems.
-
29 Feb 2024 4 repositories listedOur large model, StarCoder2- 15B, significantly outperforms other models of comparable size.
-
16 Oct 2023 4 repositories listed Syntology ran 6 of 8 samples · 2 unverifiedWe present Llemma, a large language model for mathematics.
-
18 Jul 2023 4 repositories listed Syntology ran 2 of 6 samples · 4 unverified · 5 pointer-only (licence)We find that the performance and behavior of both GPT-3.
-
6 Oct 2022 4 repositories listedFinally, we show that the multilingual reasoning abilities of language models extend to other tasks such as commonsense reasoning and word-in-context semantic judgment.
-
16 Mar 2022 4 repositories listed Syntology ran 20 of 34 samples · 14 unverified · 2 pointer-only (licence)Language models typically need to be trained or finetuned in order to acquire new knowledge, which involves updating their weights.
-
22 Apr 2025 3 repositories listed Syntology ran 7 of 16 samples · 9 unverifiedFurthermore, although TTRL is only supervised by the Maj@N metric, TTRL has demonstrated performance to consistently surpass the upper limit of the initial model, and approach the performance of models trained directly…
-
17 Mar 2025 3 repositories listed Syntology ran 1 of 5 samples · 4 unverifiedRecent breakthroughs in solving reasoning, math and coding problems with Large Language Models (LLMs) have been enabled by investing substantial computation budgets at inference time.
-
10 Feb 2025 3 repositories listedWe train our ReasonFlux-32B model with only 8 GPUs and introduces three innovations: (i) a structured and generic thought template library, containing around 500 high-level thought templates capable of generalizing to…
-
5 Feb 2025 3 repositories listedWhile conventional wisdom suggests that sophisticated reasoning tasks demand extensive training data (>100, 000 examples), we demonstrate that complex mathematical reasoning abilities can be effectively elicited with…
-
22 Jan 2025 3 repositories listed Syntology ran 3 of 10 samples · 7 unverifiedMoreover, we present effective long2short methods that use long-CoT techniques to improve short-CoT models, yielding state-of-the-art short-CoT reasoning results -- e.
-
8 Jan 2025 3 repositories listedWe present rStar-Math to demonstrate that small language models (SLMs) can rival or even surpass the math reasoning capability of OpenAI o1, without distillation from superior models.
Syntology lines on 21 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections