Home › Datasets › task › Visual Reasoning
Visual Reasoning datasets
archive 2025-07-28
44 datasets carry the task tag "Visual Reasoning" (the task itself: Visual Reasoning), ordered by the archive's paper count. Page 1 of 1: 44 shown of 44. Facet routes are this site's own (the archive records the tag string, not a page).
The archive holds 12,214 dataset rows; 12,172 are listed. 6 are withheld from every listing and count here as vandalised before snapshot (6 with contact-centre spam in the title, 0 with a spam description on a row that has no homepage, no paper and no papers counted; none with more than 1 paper, 0 with a benchmark), listed in withheld.json; 1 listed row carries a vandalised description, withheld on its page. This gate never withholds a row with a homepage or a paper that resolves, and a clean description; the content rules below withhold a row whose name is spam whatever else it carries. The gate is a phrase list: these are the rows it caught, not a claim that the rest is clean. Before that gate, the site's content rules withhold 36 more rows (invite-code, gambling, travel-booking, contact-centre and similar spam in the name or on a row with nothing real behind it); they have no page and are listed in withheld.json.
Filter 51 task tags shown of 3,717, by dataset count; the full filter by modality, task and language is on /datasets
Visual Reasoning datasets 1–44 of 44
The RefCOCO dataset is a referring expression generation (REG) dataset used for tasks related to understanding natural language expressions that refer to specific objects in images.
439 papers · 11 benchmarks
MathVista (Mathematical Reasoning of in Visual Contexts)
MathVista is a consolidated Mathematical reasoning benchmark within Visual contexts.
242 papers · 0 benchmarks
SHAPES (Swarm Heuristics based Adaptive and Penalized Estimation of Splines)
SHAPES is a dataset of synthetic images designed to benchmark systems for understanding of spatial and logical relations among multiple objects.
120 papers · 1 benchmark
Visual Entailment (VE) consists of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks.
117 papers · 2 benchmarks
Contains 145k captions for 28k images.
98 papers · 1 benchmark
RAVEN consists of 1,120,000 images and 70,000 RPM (Raven's Progressive Matrices) problems, equally distributed in 7 distinct figure configurations.
96 papers · 0 benchmarks
Winoground is a dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning.
90 papers · 1 benchmark
AbstractReasoning is a dataset for abstract reasoning, where the goal is to infer the correct answer from the context panels based on abstract reasoning.
86 papers · 0 benchmarks
NLVR (Natural Language Visual Reasoningnatural language for visual reasoning)
NLVR contains 92,244 pairs of human-written English sentences grounded in synthetic images.
83 papers · 3 benchmarks
VSR (Visual Spatial Reasoning)
The Visual Spatial Reasoning (VSR) corpus is a collection of caption-image pairs with true/false labels.
69 papers · 1 benchmark
FigureQA is a visual reasoning corpus of over one million question-answer pairs grounded in over 100,000 images.
61 papers · 1 benchmark
PGM (Procedurally Generated Matrices (PGM))
PGM dataset serves as a tool for studying both abstract reasoning and generalisation in models.
57 papers · 0 benchmarks
Rendered synthetically using a library of standard 3D objects, and tests the ability to recognize compositions of object movements that require long-term reasoning.
51 papers · 3 benchmarks
IconQA (Icon Question Answering)
Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images in the daily-life context.
42 papers · 1 benchmark
PHYRE (PHYsical REasoning)
Benchmark for physical reasoning that contains a set of simple classical mechanics puzzles in a 2D physical environment.
35 papers · 2 benchmarks
IGLUE (Image-Grounded Language Understanding Evaluation)
The Image-Grounded Language Understanding Evaluation (IGLUE) benchmark brings together—by both aggregating pre-existing datasets and creating new ones—visual question answering, cross-modal retrieval, grounded reasoning, and grounded…
31 papers · 0 benchmarks
MaRVL (Multicultural Reasoning over Vision and Language)
Multicultural Reasoning over Vision and Language (MaRVL) is a dataset based on an ImageNet-style hierarchy representative of many languages and cultures (Indonesian, Mandarin Chinese, Swahili, Tamil, and Turkish).
29 papers · 1 benchmark
Social-IQ is an unconstrained benchmark specifically designed to train and evaluate socially intelligent technologies.
23 papers · 0 benchmarks
InfiMM-Eval (Complex Open-ended Reasoning Evaluation for Multi-Modal Language Models)
Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence.
19 papers · 1 benchmark
CLEVR-Ref+ is a synthetic diagnostic dataset for referring expression comprehension.
17 papers · 1 benchmark
Super-CLEVR is a dataset for Visual Question Answering (VQA) where different factors in VQA domain shifts can be isolated in order that their effects can be studied independently.
13 papers · 0 benchmarks
Cops-Ref is a dataset for visual reasoning in context of referring expression comprehension with two main features.
8 papers · 0 benchmarks
A configurable visual question and answer dataset (COG) to parallel experiments in humans and animals.
7 papers · 0 benchmarks
The Relative Size dataset contains 486 object pairs between 41 physical objects.
6 papers · 0 benchmarks
SMART-101 (Simple Multimodal Algorithmic Reasoning Task Dataset)
Recent times have witnessed an increasing number of applications of deep neural networks towards solving tasks that require superior cognitive abilities, e.g., playing Go, generating art, ChatGPT, etc.
6 papers · 0 benchmarks
ComPhy (Compositional Physical Reasoning Dataset)
Compositional Physical Reasoning is a dataset for understanding object-centric and relational physics properties hidden from visual appearances.
5 papers · 0 benchmarks
KiloGram is a resource for studying abstract visual reasoning in humans and machines.
5 papers · 0 benchmarks
PGDP5K (Plane Geometry Diagram Parsing Dataset)
PGDP5K is a dataset consisting of 5000 diagram samples composed of 16 shapes, covering 5 positional relations, 22 symbol types and 6 text types, labeled with more fine-grained annotations at primitive level, including primitive classes,…
5 papers · 1 benchmark
TRANCE (Transformation Driven Visual Reasoning)
TRANCE extends CLEVR by asking a uniform question, i.e.
5 papers · 0 benchmarks
This dataset is collected via the WinoGAViL game to collect challenging vision-and-language associations.
5 papers · 2 benchmarks
VASR (Visual Analogies of Situation Recognition)
Visual Analogies of Situation Recognition (VASR) is a dataset for visual analogical mapping, adapting the classical word-analogy task into the visual domain.
4 papers · 1 benchmark
ADE-Affordance is a new dataset that builds upon ADE20k, which contains annotations enabling such rich visual reasoning.
3 papers · 0 benchmarks
Bongard-OpenWorld is a new benchmark for evaluating real-world few-shot reasoning for machine vision.
3 papers · 1 benchmark
A benchmark designed to evaluate MLLMs’ proficiency in understanding inter-object relationships and textual content.
3 papers · 0 benchmarks
G-VUE (General-purpose Visual Understanding Evaluation)
General-purpose Visual Understanding Evaluation (G-VUE) is a comprehensive benchmark covering the full spectrum of visual cognitive abilities with four functional domains -- Perceive, Ground, Reason, and Act.
2 papers · 0 benchmarks
The IRFL dataset consists of idioms, similes, and metaphors with matching figurative and literal images, as well as two novel tasks of multimodal figurative understanding and preference.
2 papers · 2 benchmarks
Consists of visual arithmetic problems automatically generated using a grammar model--And-Or Graph (AOG).
2 papers · 0 benchmarks
EMMA (An Enhanced MultiModal ReAsoning Benchmark)
We introduce EMMA (Enhanced MultiModal reAsoning), a benchmark targeting organic multimodal reasoning across mathematics, physics, chemistry, and coding.
1 paper · 0 benchmarks
PolyMATH, a challenging benchmark aimed at evaluating the general cognitive reasoning abilities of MLLMs.
1 paper · 0 benchmarks
SVRT (Synthetic Visual Reasoning Task)
The Synthetic Visual Reasoning Test (SVRT) is a series of 23 classification problems involving images of randomly generated shapes.
1 paper · 0 benchmarks
Sequence Consistency Evaluation (SCE) consists of a benchmark task for sequence consistency evaluation (SCE).
1 paper · 0 benchmarks
Visual Choice of Plausible Alternatives (VCOPA) is an evaluation dataset containing 380 VCOPA questions and over 1K images with various topics, which is amenable to automatic evaluation, and present the performance of baseline reasoning…
1 paper · 0 benchmarks
lilGym is a benchmark for language-conditioned reinforcement learning in visual environment based on 2,661 highly-compositional human-written natural language statements grounded in an interactive visual environment.
1 paper · 0 benchmarks
A fundamental component of human vision is our ability to parse complex visual scenes and judge the relations between their constituent objects.
0 papers · 0 benchmarks
Paper counts and descriptions are the archive's, frozen 2025-07-28; no citation counts, no stars, no trending. Sorting by "most cited" or "newest" was a live-site feature the archive does not carry.