Browse State-of-the-Art › Visual Entailment
Visual Entailment
33 papers with code · 3 benchmarks · 3 datasets archive 2025-07-28
Visual Entailment (VE) - is a task consisting of image-sentence pairs whereby a premise is defined by an image, rather than a natural language sentence as in traditional Textual Entailment tasks. The goal is to predict whether the image semantically entails the text.
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
3 leaderboard tables shown for this task, 3 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| SNLI-VE val (9 rows) | OFA | OFA: Unifying Architectures, Tasks, and Modalities Through a... | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| SNLI-VE test (8 rows) | OFA | OFA: Unifying Architectures, Tasks, and Modalities Through a... | code | Syntology ran 1 of 1 samples · 0 unverified | Compare |
| e-SNLI-VE (2 rows) | OFA-X | Harnessing the Power of Multi-Task Pretraining for Ground-Truth... | code | Syntology ran 1 of 8 samples · 7 unverified | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
3 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 33 papers with code (56 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
25 Sep 2019 7 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Different from previous work that applies joint random masking to both modalities, we use conditional masking on pre-training tasks (i.
-
4 May 2022 6 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedWe apply a contrastive loss between unimodal image and text embeddings, in addition to a captioning loss on the multimodal decoder outputs which predicts text tokens autoregressively.
-
30 Apr 2022 4 repositories listed Syntology ran 5 of 10 samples · 5 unverifiedSpatial relations are a basic part of human cognition.
-
7 Feb 2022 4 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization.
-
16 Dec 2021 4 repositories listedWe propose a cross-modal attention distillation framework to train a dual-encoder model for vision-language understanding tasks, such as visual reasoning and visual question answering.
-
13 Jul 2021 4 repositories listed Syntology ran 6 of 10 samples · 4 unverified · 9 pointer-only (licence)Most existing Vision-and-Language (V&L) models rely on pre-trained visual encoders, using a relatively small set of manually-annotated data (as compared to web-crawled data), to perceive the visual world.
-
15 Dec 2023 3 repositories listedThis work inaugurates a new domain in factual error correction for chart captions, presenting a novel evaluation mechanism, and demonstrating an effective approach to ensuring the factuality of generated chart captions.
-
7 Apr 2021 3 repositories listedAs region-based visual features usually represent parts of an image, it is challenging for existing vision-language models to fully understand the semantics from paired natural languages.
-
11 Jun 2024 2 repositories listedGrounded Multimodal Named Entity Recognition (GMNER) task aims to identify named entities, entity types and their corresponding visual regions.
-
15 Feb 2024 2 repositories listedGrounded Multimodal Named Entity Recognition (GMNER) is a nascent multimodal task that aims to identify named entities, entity types and their corresponding visual regions.
-
11 Jun 2020 2 repositories listed Syntology ran 10 of 20 samples · 10 unverified · 6 pointer-only (licence)We present VILLA, the first known effort on large-scale adversarial training for vision-and-language (V+L) representation learning.
-
19 Dec 2024 1 repository listedAdditionally, we introduce a reward-driven update optimization method to further enhance the quality of updates generated by multimodal models.
-
2 May 2024 1 repository listedTo close this gap, we propose a new task framing the figurative meaning understanding problem as an explainable visual entailment task, where the model has to predict whether the image (premise) entails a caption…
-
14 Mar 2024 1 repository listedDespite the demonstrated parameter efficiency of prompt-based multimodal fusion methods, their limited adaptivity and expressiveness often result in suboptimal performance compared to other tuning approaches.
-
5 Mar 2024 1 repository listedVisual entailment (VE) is a multimodal reasoning task consisting of image-sentence pairs whereby a promise is defined by an image, and a hypothesis is described by a sentence.
-
17 Dec 2023 1 repository listedIn this paper, we present a novel modeling framework that recasts adapter tuning after attention as a graph message passing process on attention graphs, where the projected query and value features and attention matrix…
-
4 Dec 2023 1 repository listedQVix enables a wider exploration of visual scenes, improving the LVLMs' reasoning accuracy and depth in tasks such as visual question answering and visual entailment.
-
29 Jun 2023 1 repository listed Syntology ran 0 of 4 samples · 4 unverifiedOur evaluation across three distinct tasks (image-text retrieval, visual entailment, and natural language visual reasoning) demonstrates that this approach outperforms the state-of-the-art multilingual vision-language…
-
24 May 2023 1 repository listedWe propose to solve the task through the collaboration between Large Language Models (LLMs) and Diffusion Models: Instruct GPT-3 (davinci-002) with Chain-of-Thought prompting generates text that represents a visual…
-
15 Dec 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedMultimodal image-text models have shown remarkable performance in the past few years.
-
8 Dec 2022 1 repository listed Syntology ran 1 of 8 samples · 7 unverifiedNatural language explanations promise to offer intuitively understandable explanations of a neural network's decision process in complex vision-language tasks, as pursued in recent VL-NLE models.
-
17 Nov 2022 1 repository listed Syntology ran 1 of 6 samples · 5 unverifiedWe produce models using only text training data on four representative tasks: image captioning, visual entailment, visual question answering and visual news captioning, and evaluate them on standard benchmarks using…
-
11 Oct 2022 1 repository listedMultimodal semantic understanding often has to deal with uncertainty, which means the obtained messages tend to refer to multiple targets.
-
29 Aug 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverifiedVision and Language Pretraining has become the prevalent approach for tackling multimodal downstream tasks.
-
4 Aug 2022 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Prompt tuning has become a new paradigm for model tuning and it has demonstrated success in natural language pretraining and even vision pretraining.
-
23 Jul 2022 1 repository listedCSI), a relation inferrer, and a Lexical Constraint-aware Generator (arr.
-
16 Jun 2022 1 repository listed Syntology ran 0 of 2 samples · 2 unverifiedData augmentation is a necessity to enhance data efficiency in deep learning.
-
29 Mar 2022 1 repository listedIn this paper, we propose an extension of this task, where the goal is to predict the logical relationship of fine-grained knowledge elements within a piece of text to an image.
-
9 Mar 2022 1 repository listedCurrent NLE models explain the decision-making process of a vision or vision-language model (a.
-
1 Aug 2021 1 repository listedBesides, they only explore the interaction between image and question, ignoring the semantics of candidate answers.
Syntology lines on 13 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections