Home › Datasets › task › Logical Reasoning
Logical Reasoning datasets
archive 2025-07-28
18 datasets carry the task tag "Logical Reasoning" (the task itself: Logical Reasoning), ordered by the archive's paper count. Page 1 of 1: 18 shown of 18. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Logical Reasoning datasets 1–18 of 18
BIG-bench (Beyond the Imitation Game Benchmark)
The Beyond the Imitation Game Benchmark (BIG-bench) is a collaborative benchmark intended to probe large language models and extrapolate their future capabilities.
349 papers · 121 benchmarks
By perturbing the widely used GSM8K dataset, an adversarial dataset for grade-school math called GSM-Plus is created.
17 papers · 1 benchmark
SMART-101 (Simple Multimodal Algorithmic Reasoning Task Dataset)
Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, ChatGPT, etc.
6 papers · 0 benchmarks
FLD (Formal Logic Deduction)
A deductive reasoning benchmark based on formal logic theory.
5 papers · 0 benchmarks
This dataset is a benchmark for complex reasoning abilities in large language models, drawing on United Kingdom Linguistics Olympiad problems which cover a wide range of languages.
4 papers · 1 benchmark
MultiQ is a multi-hop QA dataset for Russian, suitable for general open-domain question answering, information retrieval, and reading comprehension tasks.
3 papers · 1 benchmark
CheGeKa is a Jeopardy!-like Russian QA dataset collected from the official Russian quiz database ChGK.
2 papers · 1 benchmark
Ethics (per ethics) dataset is created to test the knowledge of the basic concepts of morality.
2 papers · 1 benchmark
Introduction Generalized quantifiers (e.g., few, most) are used to indicate the proportions predicates are satisfied.
2 papers · 0 benchmarks
RuOpenBookQA is a QA dataset with multiple-choice elementary-level science questions which probe the understanding of core science facts.
2 papers · 1 benchmark
RuWorldTree is a QA dataset with multiple-choice elementary-level science questions, which evaluate the understanding of core science facts.
2 papers · 1 benchmark
A benchmark for suppositional reasoning based on the principles of knights and knaves puzzles.
2 papers · 0 benchmarks
A collection of large languge model responses to tasks of propositional logic.
1 paper · 0 benchmarks
JustLogic is a natural language deductive reasoning dataset.
1 paper · 0 benchmarks
The ontology files, readme and statistical information can be found and browsed in the ontology library.
1 paper · 0 benchmarks
PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs.
1 paper · 0 benchmarks
REBUS (A Robust Evaluation Benchmark of Understanding Symbols)
Recent advances in large language models have led to the development of multimodal LLMs (MLLMs), which take both image data and text as an input.
1 paper · 1 benchmark
The Winograd schema challenge composes tasks with syntactic ambiguity, which can be resolved with logic and reasoning.
1 paper · 1 benchmark
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.