Papers › Incorporating Structured Representations into Pretrained Vision & Language Models...

Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene Graphs

10 May 2023arXiv:2305.06343archive 2025-07-28

Roei Herzig, Alon Mendelson, Leonid Karlinsky, Assaf Arbelle, Rogerio Feris, Trevor Darrell, Amir Globerson

Vision and language models (VLMs) have demonstrated remarkable zero-shot (ZS) performance in a variety of tasks. However, recent works have shown that even the best VLMs struggle to capture aspects of compositional scene understanding, such as object attributes, relations, and action states. In contrast, obtaining structured annotations, such as scene graphs (SGs), that could improve these models is time-consuming and costly, and thus cannot be used on a large scale. Here we ask whether small SG datasets can provide sufficient information for enhancing structured understanding of pretrained VLMs. We show that it is indeed possible to improve VLMs when learning from SGs by integrating components that incorporate structured information into both visual and textual representations. For the visual side, we incorporate a special "SG Component" in the image transformer trained to predict SG information, while for the textual side, we utilize SGs to generate fine-grained captions that highlight different compositional aspects of the scene. Our method improves the performance of several popular VLMs on multiple VL datasets with only a mild degradation in ZS capabilities.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Scene UnderstandingVisual Reasoning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Reasoning Winoground BLIP2 (SGVL) Group Score 23.3 #25 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP2 (SGVL) Image Score 28.5 #25 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP2 (SGVL) Text Score 42.8 #25 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP (SGVL) Group Score 21.5 #26 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP (SGVL) Image Score 27.3 #26 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP (SGVL) Text Score 42.8 #26 of 114 Archive leaderboard report
Visual Reasoning Winoground NegBLIP Group Score 18.5 #29 of 114 Archive leaderboard report
Visual Reasoning Winoground NegBLIP Image Score 24.0 #29 of 114 Archive leaderboard report
Visual Reasoning Winoground NegBLIP Text Score 42.5 #29 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP2 Group Score 19.0 #32 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP2 Image Score 23.8 #32 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP2 Text Score 42.0 #32 of 114 Archive leaderboard report
Visual Reasoning Winoground NegBLIP2 Group Score 20.5 #34 of 114 Archive leaderboard report
Visual Reasoning Winoground NegBLIP2 Image Score 26.0 #34 of 114 Archive leaderboard report
Visual Reasoning Winoground NegBLIP2 Text Score 41.5 #34 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP (+Graph Text, +Graph Neg) Group Score 19.0 #35 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP (+Graph Text, +Graph Neg) Image Score 25.5 #35 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP (+Graph Text, +Graph Neg) Text Score 40.5 #35 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP (+Graph Text) Group Score 16.5 #36 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP (+Graph Text) Image Score 20.5 #36 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP (+Graph Text) Text Score 40.3 #36 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP Group Score 15.0 #41 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP Image Score 19.2 #41 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP Text Score 39.0 #41 of 114 Archive leaderboard report
Visual Reasoning Winoground CLIP (SGVL) Group Score 9.8 #61 of 114 Archive leaderboard report
Visual Reasoning Winoground CLIP (SGVL) Image Score 14.0 #61 of 114 Archive leaderboard report
Visual Reasoning Winoground CLIP (SGVL) Text Score 32.0 #61 of 114 Archive leaderboard report
Visual Reasoning Winoground NegCLIP Group Score 8.0 #72 of 114 Archive leaderboard report
Visual Reasoning Winoground NegCLIP Image Score 10.5 #72 of 114 Archive leaderboard report
Visual Reasoning Winoground NegCLIP Text Score 29.5 #72 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA Group Score 13.0 #87 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA Image Score 25.0 #87 of 114 Archive leaderboard report
Visual Reasoning Winoground LLaVA Text Score 24.8 #87 of 114 Archive leaderboard report
Visual Reasoning Winoground MiniGPT-4 Group Score 9.5 #91 of 114 Archive leaderboard report
Visual Reasoning Winoground MiniGPT-4 Image Score 18.0 #91 of 114 Archive leaderboard report
Visual Reasoning Winoground MiniGPT-4 Text Score 23.3 #91 of 114 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections