{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/does-structural-attention-improve","title":"Does Structural Attention Improve Compositional Representations in Vision-Language Models?","arxiv_id":null,"date":"2022-12-03","proceeding":"NeurIPS Workshop: Self-Supervised Learning - Theory and Practice 2022 12","authors":["Rohan Pandey","Rulin Shao","Paul Pu Liang","Louis-Philippe Morency"],"abstract":"Although scaling self-supervised approaches has gained widespread success in\r\nVision-Language pre-training, a number of works providing structural knowledge of\r\nvisually-grounded semantics have recently shown incremental performance gains.\r\nPast work hypothesizes that providing structural knowledge to models in the form\r\nof scene graphs, syntax parses, etc. will result in better Structure Alignment and\r\nthus maintain representational compositionality, a core feature of human cognition.\r\nWe compare one such Structural Training model to a Structural Attention model\r\nwhich has only implicitly learned inter-modal structure alignment through a self\r\nsupervised attention regularizer. We report that the latter model results in a 52%\r\nimprovement over its baseline on the Winoground evaluation dataset, establishing\r\na new vision-language compositionality state-of-the-art (Group=16.00). We begin\r\nexploring why this self-supervised approach succeeds where a more strongly\r\nsupervised approach fails, specifically analyzing what the auxiliary loss implicitly\r\nconveys about structural knowledge.","url_abs":"https://sslneurips22.github.io/paper_pdfs/paper_65.pdf","url_pdf":"https://sslneurips22.github.io/paper_pdfs/paper_65.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"visual-reasoning","task_name":"Visual Reasoning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"IAIS large (Flickr30k)","rank_in_archive_order":30,"of":114,"metrics":{"Group Score":"16.00","Image Score":"19.75","Text Score":"42.50"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"IAIS large (COCO)","rank_in_archive_order":33,"of":114,"metrics":{"Group Score":"15.50","Image Score":"19.75","Text Score":"41.75"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"CACR base","rank_in_archive_order":38,"of":114,"metrics":{"Group Score":"14.25","Image Score":"17.75","Text Score":"39.25"},"uses_additional_data":false},{"leaderboard":"/sota/visual-reasoning-on-winoground","task":"Visual Reasoning","dataset":"Winoground","model":"ROSITA (Flickr30k)","rank_in_archive_order":52,"of":114,"metrics":{"Group Score":"12.25","Image Score":"15.25","Text Score":"35.25"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}