Papers › Going Beyond Nouns With Vision & Language Models Using Synthetic Data
Going Beyond Nouns With Vision & Language Models Using Synthetic Data
Paola Cascante-Bonilla, Khaled Shehada, James Seale Smith, Sivan Doveh, Donghyun Kim, Rameswar Panda, Gül Varol, Aude Oliva, Vicente Ordonez, Rogerio Feris, Leonid Karlinsky
Large-scale pre-trained Vision & Language (VL) models have shown remarkable performance in many applications, enabling replacing a fixed set of supported classes with zero-shot open vocabulary reasoning over (almost arbitrary) natural language prompts. However, recent works have uncovered a fundamental weakness of these models. For example, their difficulty to understand Visual Language Concepts (VLC) that go 'beyond nouns' such as the meaning of non-object words (e.g., attributes, actions, relations, states, etc.), or difficulty in performing compositional reasoning such as understanding the significance of the order of the words in a sentence. In this work, we investigate to which extent purely synthetic data could be leveraged to teach these models to overcome such shortcomings without compromising their zero-shot capabilities. We contribute Synthetic Visual Concepts (SyViC) - a million-scale synthetic dataset and data generation codebase allowing to generate additional suitable data to improve VLC understanding and compositional reasoning of VL models. Additionally, we propose a general VL finetuning strategy for effectively leveraging SyViC towards achieving these improvements. Our extensive experiments and ablations on VL-Checklist, Winoground, and ARO benchmarks demonstrate that it is possible to adapt strong pre-trained VL models with synthetic data significantly enhancing their VLC understanding (e.g. by 9.9% on ARO and 4.3% on VL-Checklist) with under 1% drop in their zero-shot accuracy.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Reasoning | Winoground | syn-CLIP | Group Score | 9.50 | #70 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | syn-CLIP | Image Score | 11.50 | #70 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | syn-CLIP | Text Score | 30.00 | #70 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | syn-CyCLIP | Group Score | 8.25 | #71 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | syn-CyCLIP | Image Score | 10.75 | #71 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | syn-CyCLIP | Text Score | 30.00 | #71 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | CyCLIP | Group Score | 7.25 | #76 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | CyCLIP | Image Score | 9.50 | #76 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | CyCLIP | Text Score | 28.50 | #76 of 114 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections