Home › Datasets › task › Mathematical Reasoning
Mathematical Reasoning datasets
archive 2025-07-28
26 datasets carry the task tag "Mathematical Reasoning" (the task itself: Mathematical Reasoning), ordered by the archive's paper count. Page 1 of 1: 26 shown of 26. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Mathematical Reasoning datasets 1–26 of 26
GSM8K is a dataset of 8.5K high quality linguistically diverse grade school math word problems created by human problem writers.
1,881 papers · 7 benchmarks
SVAMP (Simple Variations on Arithmetic Math word Problems)
A challenge set for elementary-level Math Word Problems (MWP).
362 papers · 2 benchmarks
MathVista (Mathematical Reasoning of in Visual Contexts)
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts.
242 papers · 0 benchmarks
PrOntoQA (Proof and Ontology-Generated Question-Answering)
PrOntoQA is a question-answering dataset which generates examples with chains-of-thought that describe the reasoning required to answer the questions correctly.
55 papers · 0 benchmarks
GeoQA (Geometric Question Answering)
GeoQA is a dataset for automatic geometric problem solving containing 5,010 geometric problems with corresponding annotated programs, which illustrate the solving process of the given problems Compared with another publicly available…
46 papers · 1 benchmark
IconQA (Icon Question Answering)
Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images in the daily-life context.
42 papers · 1 benchmark
We propose the first question-answering dataset driven by STEM theorems.
40 papers · 1 benchmark
MathBench (MathBench: Evaluating the Theory and Application Proficiency of LLMs with a Hierarchical Mathematics Benchmark)
MathBench is an All in One math dataset for language model evaluation, with: A Sophisticated Five-Stage Difficulty Mechanism: Unlike the usual mathematical datasets that can only evaluate a single difficulty level or have a mix of unclear…
16 papers · 0 benchmarks
A new large scale plane geometry problem solving dataset called PGPS9K, labeled both fine-grained diagram annotation and interpretable solution program.
14 papers · 1 benchmark
Lila is a unified mathematical reasoning benchmark consisting of 23 diverse tasks along four dimensions: (i) mathematical abilities e.g., arithmetic, calculus (ii) language format e.g., question-answering, fill-in-the-blanks (iii) language…
12 papers · 0 benchmarks
Math-Vision (Math-V) dataset is a meticulously curated collection of 3,040 high-quality mathematical problems with visual contexts sourced from real math competitions.
12 papers · 1 benchmark
CriticBench is a comprehensive benchmark designed to assess the abilities of Large Language Models (LLMs) to critique and rectify their reasoning across various tasks.
10 papers · 0 benchmarks
CLEVR-Math is a multi-modal math word problems dataset consisting of simple math word problems involving addition/subtraction, represented partly by a textual description and partly by an image illustrating the scenario.
8 papers · 0 benchmarks
MGSM8KInstruct, the multilingual math reasoning instruction dataset, encompassing ten distinct languages, thus addressing the issue of training data scarcity in multilingual math reasoning.
2 papers · 0 benchmarks
ParaMAWPS (Paraphrased Math Word Problem Solving Repository)
This repository contains the code, data, and models of the paper titled "Math Word Problem Solving by Generating Linguistic Variants of Problem Statements" published in the Proceedings of the 61st Annual Meeting of the Association for…
2 papers · 1 benchmark
QUITE (Quantifying Uncertainty in natural language Text)
QUITE (Quantifying Uncertainty in natural language Text) is an entirely new benchmark that allows for assessing the capabilities of neural language model-based systems w.r.t.
2 papers · 0 benchmarks
ASyMOB (ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark)
ASyMOB (pronounced Asimov, in tribute to the renowned author), is a novel assessment framework focused exclusively on symbolic manipulation, featuring 17,092 unique math challenges, organized by similarity and complexity.
1 paper · 0 benchmarks
CHAMP (Concept and Hint-Annotated Math Problems)
The Concept and Hint-Annotated Math Problems (CHAMP) consists of high school math competition problems, annotated with concepts, or general math facts, and hints, or problem-specific tricks.
1 paper · 0 benchmarks
Conic10K is an open-ended math problem dataset on conic sections in Chinese senior high school education.
1 paper · 0 benchmarks
Enumerate–Conjecture–Prove: Formally Solving Answer-Construction Problem in Math Competitions We release the ConstructiveBench dataset as part of our Enumerate–Conjecture–Prove (ECP) paper.
1 paper · 0 benchmarks
Existing arithmetic benchmarks have a limited number of multiple-choice questions.
1 paper · 1 benchmark
Existing arithmetic benchmarks have a limited number of True-or-False questions.
1 paper · 1 benchmark
PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs.
1 paper · 0 benchmarks
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
1 paper · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.