Datasets › Winoground

Winoground

Introduced by Tristan Thrush et al. in Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality7 Apr 2022 archive 2025-07-28

Winoground is a dataset for evaluating the ability of vision and language models to conduct visio-linguistic compositional reasoning. Given two images and two captions, the goal is to match them correctly -- but crucially, both captions contain a completely identical set of words, only in a different order. The dataset was carefully hand-curated by expert annotators and is labeled with a rich set of fine-grained tags to assist in analyzing model performance.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Reasoning Winoground GPT-4o + CA Text Score 75.5 A Cognitive Paradigm Approach to Probe the... — 114 Compare

Papers archive 2025-07-28

19 shown of 19 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 90. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
A Cognitive Paradigm Approach to Probe the Perception-Reasoning Interface in VLMs 0 1 23 Jan 2025 not harvested
Prompting Large Vision-Language Models for Compositional Reasoning 1 3 20 Jan 2024 ran 1 of 2 samples (1 unverified; 2 pointer-only for licence)
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs 1 13 5 Jan 2024 not harvested
Compositional Chain-of-Thought Prompting for Large Multimodal Models 1 6 27 Nov 2023 ran 3 of 4 samples (1 unverified)
SelfEval: Leveraging the discriminative nature of generative models for evaluation 0 6 17 Nov 2023 not harvested
The Role of Chain-of-Thought in Complex Vision-Language Reasoning Task 0 2 15 Nov 2023 not harvested
MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning 2 1 14 Sep 2023 ran 5 of 7 samples (2 unverified; 7 pointer-only for licence)
An Examination of the Compositionality of Large Generative Vision-Language Models 1 5 21 Aug 2023 not harvested
Revisiting the Role of Language Priors in Vision-Language Models 1 3 2 Jun 2023 not harvested
What You See is What You Read? Improving Text-Image Alignment Evaluation 1 8 17 May 2023 not harvested
Measuring Progress in Fine-grained Vision-and-Language Understanding 2 9 12 May 2023 ran 2 of 4 samples (2 unverified; 1 pointer-only for licence)
Simple Token-Level Confidence Improves Caption Correctness 0 6 11 May 2023 not harvested
Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene Graphs 0 12 10 May 2023 not harvested
Going Beyond Nouns With Vision & Language Models Using Synthetic Data 1 3 30 Mar 2023 not harvested
Your Diffusion Model is Secretly a Zero-Shot Classifier 4 1 28 Mar 2023 ran 2 of 2 samples (0 unverified; 1 pointer-only for licence)
Equivariant Similarity for Vision-Language Foundation Models 1 6 25 Mar 2023 ran 1 of 1 samples (0 unverified)
ViLEM: Visual-Language Error Modeling for Image-Text Retrieval 0 2 1 Jan 2023 not harvested
Does Structural Attention Improve Compositional Representations in Vision-Language Models? 0 4 3 Dec 2022 not harvested
Winoground: Probing Vision and Language Models for Visio-Linguistic Compositionality 2 20 7 Apr 2022 ran 3 of 4 samples (1 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Winoground

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections