Datasets › CLEVR

CLEVR (Compositional Language and Elementary Visual Reasoning)

Introduced by Justin Johnson et al. in CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning20 Dec 2016 archive 2025-07-28

CLEVR (Compositional Language and Elementary Visual Reasoning) is a synthetic Visual Question Answering dataset. It contains images of 3D-rendered objects; each image comes with a number of highly compositional questions that fall into different categories. Those categories fall into 5 classes of tasks: Exist, Count, Compare Integer, Query Attribute and Compare Attribute. The CLEVR dataset consists of: a training set of 70k images and 700k questions, a validation set of 15k images and 150k questions, a test set of 15k images and 150k questions about objects, answers, scene graphs and functional programs for all train and validation images and questions. Each object present in the scene, aside of position, is characterized by a set of four attributes: 2 sizes: large, small, 3 shapes: square, cylinder, sphere, 2 material types: rubber, metal, 8 color types: gray, blue, brown, yellow, red, green, purple, cyan, resulting in 96 unique combinations.

Source: On transfer learning using a MAC model variant Image Source: Johnson et al

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Question Answering (VQA) CLEVR NS-VQA (1K programs) Accuracy 99.8 Neural-Symbolic VQA: Disentangling Reasoning from Vision... kexinyi/ns-vqa +1 15 Compare
Image Generation CLEVR Projected GAN FID-5k-training-steps 0.89 Projected GANs Converge Faster autonomousvision/projected_gan +2 6 Compare
Visual Question Answering CLEVR NeSyCoCo Neuro-Symbolic Accuracy 99.7 NeSyCoCo: A Neuro-Symbolic Concept Composer for... hlr/nesycoco 1 Compare

Papers archive 2025-07-28

17 shown of 17 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 657. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
NeSyCoCo: A Neuro-Symbolic Concept Composer for Compositional Generalization 1 2 20 Dec 2024 not harvested
Projected GANs Converge Faster 3 1 1 Nov 2021 ran 38 of 49 samples (11 unverified; 6 pointer-only for licence)
MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding 5 1 26 Apr 2021 ran 6 of 11 samples (5 unverified)
Generative Adversarial Transformers 2 5 1 Mar 2021 ran 5 of 6 samples (1 unverified)
Interpretable Visual Reasoning via Induced Symbolic Space 1 1 23 Nov 2020 not harvested
Language-Conditioned Graph Networks for Relational Reasoning 1 1 10 May 2019 ran 0 of 10 samples (10 unverified)
The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision 2 1 26 Apr 2019 ran 0 of 7 samples (7 unverified)
Explainable and Explicit Visual Reasoning over Scene Graphs 2 1 5 Dec 2018 ran 2 of 3 samples (1 unverified)
Neural-Symbolic VQA: Disentangling Reasoning from Vision and Language Understanding 2 1 4 Oct 2018 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Question-Guided Hybrid Convolution for Visual Question Answering 0 1 8 Aug 2018 not harvested
Learning Visual Question Answering by Bootstrapping Hard Attention 1 1 1 Aug 2018 not harvested
DDRprog: A CLEVR Differentiable Dynamic Reasoning Programmer 0 1 30 Mar 2018 not harvested
Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning 1 1 14 Mar 2018 ran 0 of 8 samples (8 unverified)
Compositional Attention Networks for Machine Reasoning 10 1 8 Mar 2018 ran 1 of 7 samples (6 unverified)
FiLM: Visual Reasoning with a General Conditioning Layer 7 1 22 Sep 2017 ran 9 of 10 samples (1 unverified; 7 pointer-only for licence)
A simple neural network module for relational reasoning 20 1 5 Jun 2017 ran 3 of 7 samples (4 unverified; 1 pointer-only for licence)
Inferring and Executing Programs for Visual Reasoning 5 1 10 May 2017 ran 4 of 6 samples (2 unverified; 6 pointer-only for licence)

Dataset loaders archive 2025-07-28

3 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • CLEVR
  • CLEVR-CoGenT

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections