Browse State-of-the-Art › Visual Commonsense Reasoning
Visual Commonsense Reasoning
33 papers with code · 7 benchmarks · 8 datasets archive 2025-07-28
Image source: Visual Commonsense Reasoning
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
7 leaderboard tables shown for this task, 7 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
8 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
30 shown of 33 papers with code (65 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
6 Aug 2019 11 repositories listed Syntology ran 10 of 34 samples · 24 unverified · 34 pointer-only (licence)We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language.
-
25 Sep 2019 7 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Different from previous work that applies joint random masking to both modalities, we use conditional masking on pre-training tasks (i.
-
1 Dec 2023 4 repositories listedFurthermore, we present ViP-Bench, a comprehensive benchmark to assess the capability of models in understanding visual prompts across multiple dimensions, enabling future research in this domain.
-
27 Nov 2018 4 repositories listed Syntology ran 0 of 6 samples · 6 unverifiedWhile this task is easy for humans, it is tremendously difficult for today's vision systems, requiring higher-order cognition and commonsense reasoning about the world.
-
7 Jul 2023 3 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 5 pointer-only (licence)Before sending to LLM, the reference is replaced by RoI features and interleaved with language embeddings as a sequence.
-
22 Aug 2019 3 repositories listed Syntology ran 0 of 1 samples · 1 unverifiedWe introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short).
-
18 Aug 2021 2 repositories listedNevertheless, there has not been an open-source codebase in support of training and deploying numerous neural network models for cross-modal analytics in a unified and modular fashion.
-
4 Feb 2021 2 repositories listed Syntology ran 2 of 12 samples · 10 unverifiedOn 7 popular vision-and-language benchmarks, including visual question answering, referring expression comprehension, visual commonsense reasoning, most of which have been previously modeled as discriminative tasks, our…
-
11 Jun 2020 2 repositories listed Syntology ran 10 of 20 samples · 10 unverified · 6 pointer-only (licence)We present VILLA, the first known effort on large-scale adversarial training for vision-and-language (V+L) representation learning.
-
8 Dec 2024 1 repository listedLarge multimodal models (LMMs) have shown remarkable performance in the visual commonsense reasoning (VCR) task, which aims to answer a multiple-choice question based on visual commonsense within an image.
-
19 Jun 2024 1 repository listedTo facilitate multimodal grounded language modeling, we employ a late-fusion layer that combines the projected visual features with the output of a pre-trained LLM conditioned on text only.
-
3 Jun 2024 1 repository listed Syntology ran 1 of 1 samples · 0 unverified · 1 pointer-only (licence)Recent advances in vision-language models (VLMs) have demonstrated the advantages of processing images at higher resolutions and utilizing multi-crop features to preserve native resolution details.
-
5 Sep 2023 1 repository listedIn recent years, cross-modal reasoning (CMR), the process of understanding and reasoning across different modalities, has emerged as a pivotal area with applications spanning from multimedia analysis to healthcare…
-
1 Jan 2023 1 repository listedLanguage models are capable of commonsense reasoning: while domain-specific models can learn from explicit knowledge (e.
-
8 Dec 2022 1 repository listed Syntology ran 0 of 8 samples · 8 unverifiedWe leverage situation recognition annotations and the CLIP model to generate a large set of 500k candidate analogies.
-
20 Sep 2022 1 repository listed Syntology ran 4 of 5 samples · 1 unverified · 2 pointer-only (licence)We further design language models to learn to generate lectures and explanations as the chain of thought (CoT) to mimic the multi-hop reasoning process when answering ScienceQA questions.
-
17 Aug 2022 1 repository listedBootstrapping from pre-trained language models has been proven to be an efficient approach for building vision-language models (VLM) for tasks such as image captioning or visual question answering.
-
23 May 2022 1 repository listed Syntology ran 3 of 7 samples · 4 unverifiedWe show that PEVL enables state-of-the-art performance of detector-free VLP models on position-sensitive tasks such as referring expression comprehension and phrase grounding, and also improves the performance on…
-
30 Mar 2022 1 repository listed Syntology ran 0 of 3 samples · 3 unverifiedBreakthroughs in transformer-based models have revolutionized not only the NLP field, but also vision and multimodal systems.
-
14 Mar 2022 1 repository listedIn this work, we for the first time introduce an end-to-end video-language model, namely \textit{all-in-one Transformer}, that embeds raw video and textual signals into joint representations using a unified backbone…
-
25 Feb 2022 1 repository listedGiven that our framework is model-agnostic, we apply it to the existing popular baselines and validate its effectiveness on the benchmark dataset.
-
27 Oct 2021 1 repository listedTo overcome this limitation and take a solid step towards artificial general intelligence (AGI), we develop a foundation model pre-trained with huge multimodal data, which can be quickly adapted for various downstream…
-
14 Sep 2021 1 repository listed Syntology ran 0 of 9 samples · 9 unverifiedCommonsense is defined as the knowledge that is shared by everyone.
-
6 Aug 2021 1 repository listedWhile image understanding on recognition-level has achieved remarkable advancements, reliable visual scene understanding requires comprehensive image understanding on recognition-level but also cognition-level, which…
-
4 Jul 2021 1 repository listedMoreover, the proposed model provides intuitive interpretation into visual commonsense reasoning.
-
4 Jun 2021 1 repository listed Syntology ran 0 of 14 samples · 14 unverifiedAs humans, we understand events in the visual world contextually, performing multimodal reasoning across time to make inferences about the past, present, and future.
-
15 Oct 2020 1 repository listedNatural language rationales could provide intuitive, higher-level explanations that are easily understandable by humans, complementing the more broadly studied lower-level explanations based on gradients or attention…
-
1 Dec 2019 1 repository listedInspired by this idea, towards VCR, we propose a connective cognition network (CCN) to dynamically reorganize the visual neuron connectivity that is contextualized by the meaning of questions and answers.
-
1 Dec 2019 1 repository listedDespite impressive recent progress that has been reported on tasks that necessitate reasoning, such as visual question answering and visual dialog, models often exploit biases in datasets.
-
31 Oct 2019 1 repository listedDespite impressive recent progress that has been reported on tasks that necessitate reasoning, such as visual question answering and visual dialog, models often exploit biases in datasets.
Syntology lines on 14 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections