Datasets › NLVR

NLVR (Natural Language Visual Reasoningnatural language for visual reasoning)

Introduced by Alane Suhr et al. in A Corpus of Natural Language for Visual Reasoning1 Jan 2017 archive 2025-07-28

NLVR contains 92,244 pairs of human-written English sentences grounded in synthetic images. Because the images are synthetically generated, this dataset can be used for semantic parsing.

Source: http://lil.nlp.cornell.edu/nlvr/ Image Source: http://lil.nlp.cornell.edu/nlvr/

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

16 shown of 16 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 83. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Implicit Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis 1 1 21 Sep 2023 not harvested
Differentiable Outlier Detection Enable Robust Deep Multimodal Analysis 1 1 11 Feb 2023 not harvested
Toward Building General Foundation Models for Language, Vision, and Vision-Language Understanding Tasks 1 2 12 Jan 2023 not harvested
X²-VLM: All-In-One Pre-trained Model For Vision-Language Tasks 2 4 22 Nov 2022 ran 2 of 6 samples (4 unverified; 6 pointer-only for licence)
Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks 2 2 22 Aug 2022 not harvested
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 2 4 May 2022 ran 9 of 17 samples (8 unverified)
BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation 9 1 28 Jan 2022 not harvested
Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts 1 2 16 Nov 2021 ran 1 of 1 samples (0 unverified)
VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts 2 2 3 Nov 2021 not harvested
SimVLM: Simple Visual Language Model Pretraining with Weak Supervision 2 2 24 Aug 2021 ran 18 of 37 samples (19 unverified; 28 pointer-only for licence)
Align before Fuse: Vision and Language Representation Learning with Momentum Distillation 6 2 16 Jul 2021 ran 3 of 5 samples (2 unverified; 3 pointer-only for licence)
Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning 3 2 7 Apr 2021 not harvested
ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision 6 2 5 Feb 2021 ran 1 of 4 samples (3 unverified; 1 pointer-only for licence)
UNITER: UNiversal Image-TExt Representation Learning 7 1 25 Sep 2019 ran 3 of 3 samples (0 unverified; 2 pointer-only for licence)
LXMERT: Learning Cross-Modality Encoder Representations from Transformers 9 2 20 Aug 2019 ran 4 of 15 samples (11 unverified; 3 pointer-only for licence)
VisualBERT: A Simple and Performant Baseline for Vision and Language 10 2 9 Aug 2019 ran 4 of 9 samples (5 unverified; 6 pointer-only for licence)

Dataset loaders archive 2025-07-28

1 loader as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

Unknown

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • NLVR
  • NLVR2 Dev
  • NLVR2 Test

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections