Papers › Going Beyond Nouns With Vision & Language Models Using Synthetic Data

Going Beyond Nouns With Vision & Language Models Using Synthetic Data

30 Mar 2023ICCV 2023 1arXiv:2303.17590archive 2025-07-28

Paola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh, Donghyun Kim, Rameswar Panda, Gül Varol, Aude Oliva, Vicente Ordonez, Rogerio Feris, Leonid Karlinsky

Large-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot open vocabulary reasoning over (almost arbitrary) natural language prompts. However, recent works have uncovered a fundamental weakness of these models. For example, their difficulty to understand Visual Language Concepts (VLC) that go 'beyond nouns' such as the meaning of non-object words (e.g., attributes, actions, relations, states, etc.), or difficulty in performing compositional reasoning such as understanding the significance of the order of the words in a sentence. In this work, we investigate to which extent purely synthetic data could be leveraged to teach these models to overcome such shortcomings without compromising their zero-shot capabilities. We contribute Synthetic Visual Concepts (SyViC) - a million-scale synthetic dataset and data generation codebase allowing to generate additional suitable data to improve VLC understanding and compositional reasoning of VL models. Additionally, we propose a general VL finetuning strategy for effectively leveraging SyViC towards achieving these improvements. Our extensive experiments and ablations on VL-Checklist, Winoground, and ARO benchmarks demonstrate that it is possible to adapt strong pre-trained VL models with synthetic data significantly enhancing their VLC understanding (e.g. by 9.9% on ARO and 4.3% on VL-Checklist) with under 1% drop in their zero-shot accuracy.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

uvavision/syvic officialmentioned in paperpytorch report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

SentenceVisual Reasoning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Reasoning Winoground syn-CLIP Group Score 9.50 #70 of 114 Archive leaderboard report
Visual Reasoning Winoground syn-CLIP Image Score 11.50 #70 of 114 Archive leaderboard report
Visual Reasoning Winoground syn-CLIP Text Score 30.00 #70 of 114 Archive leaderboard report
Visual Reasoning Winoground syn-CyCLIP Group Score 8.25 #71 of 114 Archive leaderboard report
Visual Reasoning Winoground syn-CyCLIP Image Score 10.75 #71 of 114 Archive leaderboard report
Visual Reasoning Winoground syn-CyCLIP Text Score 30.00 #71 of 114 Archive leaderboard report
Visual Reasoning Winoground CyCLIP Group Score 7.25 #76 of 114 Archive leaderboard report
Visual Reasoning Winoground CyCLIP Image Score 9.50 #76 of 114 Archive leaderboard report
Visual Reasoning Winoground CyCLIP Text Score 28.50 #76 of 114 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections