Browse State-of-the-Art › Visual Reasoning
Visual Reasoning
356 papers with code · 12 benchmarks · 44 datasets archive 2025-07-28
Ability to understand actions and reasoning associated with any visual images
Description from the archive archive 2025-07-28.
Benchmarks archive 2025-07-28
12 leaderboard tables shown for this task, 12 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted. 10 shown of 12 until expanded.
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
44 datasets whose archive record lists this task, ordered by the archive's paper count. 30 shown of 44 until expanded.
Subtasks archive 2025-07-28
1 subtask in the archive's task tree.
Most implemented papers archive 2025-07-28
30 shown of 356 papers with code (698 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
26 Feb 2021 82 repositories listed Syntology ran 16 of 20 samples · 4 unverified · 16 pointer-only (licence)State-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories.
-
30 Jan 2023 17 repositories listed Syntology ran 4 of 8 samples · 4 unverified · 1 pointer-only (licence)The cost of vision-and-language pre-training has become increasingly prohibitive due to end-to-end training of large-scale models.
-
17 Apr 2023 13 repositories listed Syntology ran 16 of 51 samples · 35 unverifiedInstruction tuning large language models (LLMs) using machine-generated instruction-following data has improved zero-shot capabilities on new tasks, but the idea is less explored in the multimodal field.
-
6 Aug 2019 11 repositories listed Syntology ran 10 of 34 samples · 24 unverified · 34 pointer-only (licence)We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language.
-
9 Aug 2019 10 repositories listed Syntology ran 4 of 9 samples · 5 unverified · 6 pointer-only (licence)We propose VisualBERT, a simple and flexible framework for modeling a broad range of vision-and-language tasks.
-
8 Mar 2018 10 repositories listed Syntology ran 1 of 7 samples · 6 unverifiedWe present the MAC network, a novel fully differentiable neural network architecture, designed to facilitate explicit and expressive reasoning.
-
18 Jul 2017 10 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 4 pointer-only (licence)We present a new technique for learning visual-semantic embeddings for cross-modal retrieval.
-
28 Jan 2022 9 repositories listedFurthermore, performance improvement has been largely achieved by scaling up the dataset with noisy image-text pairs collected from the web, which is a suboptimal source of supervision.
-
20 Aug 2019 9 repositories listed Syntology ran 4 of 15 samples · 11 unverified · 3 pointer-only (licence)In LXMERT, we build a large-scale Transformer model that consists of three encoders: an object relationship encoder, a language encoder, and a cross-modality encoder.
-
2 Jan 2021 7 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)In our experiments we feed the visual features generated by the new object detection model into a Transformer-based VL fusion model \oscar \cite{li2020oscar}, and utilize an improved approach \short\ to pre-train the VL…
-
25 Sep 2019 7 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 2 pointer-only (licence)Different from previous work that applies joint random masking to both modalities, we use conditional masking on pre-training tasks (i.
-
22 Sep 2017 7 repositories listed Syntology ran 9 of 10 samples · 1 unverified · 7 pointer-only (licence)We introduce a general-purpose conditioning method for neural networks called FiLM: Feature-wise Linear Modulation.
-
20 Apr 2023 6 repositories listedOur work, for the first time, uncovers that properly aligning the visual features with an advanced large language model can possess numerous advanced multi-modal abilities demonstrated by GPT-4, such as detailed image…
-
4 May 2022 6 repositories listed Syntology ran 9 of 17 samples · 8 unverifiedWe apply a contrastive loss between unimodal image and text embeddings, in addition to a captioning loss on the multimodal decoder outputs which predicts text tokens autoregressively.
-
16 Jul 2021 6 repositories listed Syntology ran 3 of 5 samples · 2 unverified · 3 pointer-only (licence)Most existing methods employ a transformer-based multimodal encoder to jointly model visual tokens (region-based image features) and word tokens.
-
5 Feb 2021 6 repositories listed Syntology ran 1 of 4 samples · 3 unverified · 1 pointer-only (licence)Vision-and-Language Pre-training (VLP) has improved performance on various joint vision-and-language downstream tasks.
-
5 Dec 2018 6 repositories listed Syntology ran 4 of 14 samples · 10 unverifiedWe propose to compose dynamic tree structures that place the objects in an image into a visual context, helping visual reasoning tasks such as scene graph generation and visual Q&A.
-
27 Nov 2023 5 repositories listed Syntology ran 6 of 14 samples · 8 unverified · 1 pointer-only (licence)We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning.
-
25 Feb 2019 5 repositories listedWe introduce GQA, a new dataset for real-world visual reasoning and compositional question answering, seeking to address key shortcomings of previous VQA datasets.
-
10 May 2017 5 repositories listed Syntology ran 4 of 6 samples · 2 unverified · 6 pointer-only (licence)Existing methods for visual reasoning attempt to directly map inputs to outputs using black-box architectures without explicitly modeling the underlying reasoning processes.
-
20 Dec 2016 5 repositories listed Syntology ran 2 of 5 samples · 3 unverified · 4 pointer-only (licence)When building artificial intelligence systems that can reason and answer questions about visual data, we need diagnostic tests to analyze our progress and discover shortcomings.
-
18 Jul 2023 4 repositories listed Syntology ran 2 of 6 samples · 4 unverified · 5 pointer-only (licence)We find that the performance and behavior of both GPT-3.
-
28 Mar 2023 4 repositories listed Syntology ran 2 of 2 samples · 0 unverified · 1 pointer-only (licence)Our generative approach to classification, which we call Diffusion Classifier, attains strong results on a variety of benchmarks and outperforms alternative methods of extracting knowledge from diffusion models.
-
30 Apr 2022 4 repositories listed Syntology ran 5 of 10 samples · 5 unverifiedSpatial relations are a basic part of human cognition.
-
7 Feb 2022 4 repositories listed Syntology ran 1 of 1 samples · 0 unverifiedIn this work, we pursue a unified paradigm for multimodal pretraining to break the scaffolds of complex task/modality-specific customization.
-
16 Dec 2021 4 repositories listedWe propose a cross-modal attention distillation framework to train a dual-encoder model for vision-language understanding tasks, such as visual reasoning and visual question answering.
-
8 Dec 2021 4 repositories listedState-of-the-art vision and vision-and-language models rely on large-scale visio-linguistic pretraining for obtaining good performance on a variety of downstream tasks.
-
9 Jul 2019 4 repositories listed Syntology ran 3 of 21 samples · 18 unverified · 2 pointer-only (licence)We introduce the Neural State Machine, seeking to bridge the gap between the neural and symbolic views of AI and integrate their complementary strengths for the task of visual reasoning.
-
15 Apr 2024 3 repositories listed Syntology ran 13 of 16 samples · 3 unverified · 16 pointer-only (licence)Programming often involves converting detailed and complex specifications into code, a process during which developers typically utilize visual aids to more effectively convey concepts.
-
30 Mar 2022 3 repositories listed Syntology ran 5 of 7 samples · 2 unverifiedTo implement this idea, we propose Collaborative Glance-Gaze TransFormer (CoFormer) that consists of two modules: Glance transformer for activity classification and Gaze transformer for entity estimation.
Syntology lines on 25 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections