Papers › Does Structural Attention Improve Compositional Representations in Vision-Language Models?

Does Structural Attention Improve Compositional Representations in Vision-Language Models?

3 Dec 2022NeurIPS Workshop: Self-Supervised Learning - Theory and Practice 2022 12archive 2025-07-28

Rohan Pandey, Rulin Shao, Paul Pu Liang, Louis-Philippe Morency

Although scaling self-supervised approaches has gained widespread success in Vision-Language pre-training, a number of works providing structural knowledge of visually-grounded semantics have recently shown incremental performance gains. Past work hypothesizes that providing structural knowledge to models in the form of scene graphs, syntax parses, etc. will result in better Structure Alignment and thus maintain representational compositionality, a core feature of human cognition. We compare one such Structural Training model to a Structural Attention model which has only implicitly learned inter-modal structure alignment through a self supervised attention regularizer. We report that the latter model results in a 52% improvement over its baseline on the Winoground evaluation dataset, establishing a new vision-language compositionality state-of-the-art (Group=16.00). We begin exploring why this self-supervised approach succeeds where a more strongly supervised approach fails, specifically analyzing what the auxiliary loss implicitly conveys about structural knowledge.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Visual Reasoning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Reasoning Winoground IAIS large (Flickr30k) Group Score 16.00 #30 of 114 Archive leaderboard report
Visual Reasoning Winoground IAIS large (Flickr30k) Image Score 19.75 #30 of 114 Archive leaderboard report
Visual Reasoning Winoground IAIS large (Flickr30k) Text Score 42.50 #30 of 114 Archive leaderboard report
Visual Reasoning Winoground IAIS large (COCO) Group Score 15.50 #33 of 114 Archive leaderboard report
Visual Reasoning Winoground IAIS large (COCO) Image Score 19.75 #33 of 114 Archive leaderboard report
Visual Reasoning Winoground IAIS large (COCO) Text Score 41.75 #33 of 114 Archive leaderboard report
Visual Reasoning Winoground CACR base Group Score 14.25 #38 of 114 Archive leaderboard report
Visual Reasoning Winoground CACR base Image Score 17.75 #38 of 114 Archive leaderboard report
Visual Reasoning Winoground CACR base Text Score 39.25 #38 of 114 Archive leaderboard report
Visual Reasoning Winoground ROSITA (Flickr30k) Group Score 12.25 #52 of 114 Archive leaderboard report
Visual Reasoning Winoground ROSITA (Flickr30k) Image Score 15.25 #52 of 114 Archive leaderboard report
Visual Reasoning Winoground ROSITA (Flickr30k) Text Score 35.25 #52 of 114 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections