Papers › Does Structural Attention Improve Compositional Representations in Vision-Language Models?
Does Structural Attention Improve Compositional Representations in Vision-Language Models?
Rohan Pandey, Rulin Shao, Paul Pu Liang, Louis-Philippe Morency
Although scaling self-supervised approaches has gained widespread success in Vision-Language pre-training, a number of works providing structural knowledge of visually-grounded semantics have recently shown incremental performance gains. Past work hypothesizes that providing structural knowledge to models in the form of scene graphs, syntax parses, etc. will result in better Structure Alignment and thus maintain representational compositionality, a core feature of human cognition. We compare one such Structural Training model to a Structural Attention model which has only implicitly learned inter-modal structure alignment through a self supervised attention regularizer. We report that the latter model results in a 52% improvement over its baseline on the Winoground evaluation dataset, establishing a new vision-language compositionality state-of-the-art (Group=16.00). We begin exploring why this self-supervised approach succeeds where a more strongly supervised approach fails, specifically analyzing what the auxiliary loss implicitly conveys about structural knowledge.
Code
No code repository is listed for this paper in the archive or in Syntology's graph.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Reasoning | Winoground | IAIS large (Flickr30k) | Group Score | 16.00 | #30 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | IAIS large (Flickr30k) | Image Score | 19.75 | #30 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | IAIS large (Flickr30k) | Text Score | 42.50 | #30 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | IAIS large (COCO) | Group Score | 15.50 | #33 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | IAIS large (COCO) | Image Score | 19.75 | #33 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | IAIS large (COCO) | Text Score | 41.75 | #33 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | CACR base | Group Score | 14.25 | #38 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | CACR base | Image Score | 17.75 | #38 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | CACR base | Text Score | 39.25 | #38 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | ROSITA (Flickr30k) | Group Score | 12.25 | #52 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | ROSITA (Flickr30k) | Image Score | 15.25 | #52 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | ROSITA (Flickr30k) | Text Score | 35.25 | #52 of 114 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections