Datasets › SNLI-VE

SNLI-VE

Introduced by Ning Xie et al. in Visual Entailment: A Novel Task for Fine-Grained Image Understanding archive 2025-07-28

Visual Entailment (VE) consists of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks. The goal of a trained VE model is to predict whether the image semantically entails the text. SNLI-VE is a dataset for VE which is based on the Stanford Natural Language Inference corpus and Flickr30k dataset.

Source: https://github.com/necla-ml/SNLI-VE

Benchmarks archive 2025-07-28

All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Visual Entailment SNLI-VE val OFA Accuracy 91.0 OFA: Unifying Architectures, Tasks, and Modalities... modelscope/modelscope +3 9 Compare
Visual Entailment SNLI-VE test OFA Accuracy 91.2 OFA: Unifying Architectures, Tasks, and Modalities... modelscope/modelscope +3 8 Compare

Papers archive 2025-07-28

10 shown of 10 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 117. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Prompt Tuning for Generative Multimodal Pretrained Models 1 2 4 Aug 2022 ran 1 of 1 samples (0 unverified; 1 pointer-only for licence)
CoCa: Contrastive Captioners are Image-Text Foundation Models 6 2 4 May 2022 ran 9 of 17 samples (8 unverified)
Multimodal Adaptive Distillation for Leveraging Unimodal Encoders for Vision-Language Tasks 0 1 22 Apr 2022 not harvested
OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework 4 2 7 Feb 2022 ran 1 of 1 samples (0 unverified)
SimVLM: Simple Visual Language Model Pretraining with Weak Supervision 2 2 24 Aug 2021 ran 18 of 37 samples (19 unverified; 28 pointer-only for licence)
How Much Can CLIP Benefit Vision-and-Language Tasks? 4 1 13 Jul 2021 ran 6 of 10 samples (4 unverified; 9 pointer-only for licence)
Seeing Out of tHe bOx: End-to-End Pre-training for Vision-Language Representation Learning 3 2 7 Apr 2021 not harvested
Large-Scale Adversarial Training for Vision-and-Language Representation Learning 2 1 11 Jun 2020 ran 10 of 20 samples (10 unverified; 6 pointer-only for licence)
UNITER: UNiversal Image-TExt Representation Learning 7 2 25 Sep 2019 ran 3 of 3 samples (0 unverified; 2 pointer-only for licence)
Visual Entailment: A Novel Task for Fine-Grained Image Understanding 1 2 20 Jan 2019 not harvested

Dataset loaders archive 2025-07-28

2 loaders as listed in the archive; links are outbound and not re-checked here.

Tasks archive 2025-07-28

License archive 2025-07-28

No licence recorded in the archive. Absence here is not a statement about the dataset's terms.

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • SNLI-VE val
  • SNLI-VE test
  • SNLI-VE

3 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections