Home › Datasets › task › Arithmetic Reasoning

Arithmetic Reasoning datasets

archive 2025-07-28

8 datasets carry the task tag "Arithmetic Reasoning" (the task itself: Arithmetic Reasoning), ordered by the archive's paper count. Page 1 of 1: 8 shown of 8. Facet routes are this site's own (the archive records the tag string, not a page).

The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.

Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets

Arithmetic Reasoning datasets 1–8 of 8

GSM8K is a dataset of 8.5K high quality linguistically diverse grade school math word problems created by human problem writers.
1,881 papers · 7 benchmarks
MGSM (Multilingual Grade School Math)
Multilingual Grade School Math Benchmark (MGSM) is a benchmark of grade-school math problems.
107 papers · 1 benchmark
Game of 24 is a mathematical reasoning challenge, where the goal is to use 4 numbers and basic arithmetic operations (+-/) to obtain 24.
61 papers · 1 benchmark
By perturbing the widely used GSM8K dataset, an adversarial dataset for grade-school math called GSM-Plus is created.
17 papers · 1 benchmark
SMART-101 (Simple Multimodal Algorithmic Reasoning Task Dataset)
Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, ChatGPT, etc.
6 papers · 0 benchmarks
Existing arithmetic benchmarks have a limited number of multiple-choice questions.
1 paper · 1 benchmark
Existing arithmetic benchmarks have a limited number of True-or-False questions.
1 paper · 1 benchmark
PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs.
1 paper · 0 benchmarks

Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.