Papers › Revisiting the Role of Language Priors in Vision-Language Models
Revisiting the Role of Language Priors in Vision-Language Models
Zhiqiu Lin, Xinyue Chen, Deepak Pathak, Pengchuan Zhang, Deva Ramanan
Vision-language models (VLMs) are impactful in part because they can be applied to a variety of visual understanding tasks in a zero-shot fashion, without any fine-tuning. We study generative VLMs that are trained for next-word generation given an image. We explore their zero-shot performance on the illustrative task of image-text retrieval across 8 popular vision-language benchmarks. Our first observation is that they can be repurposed for discriminative tasks (such as image-text retrieval) by simply computing the match score of generating a particular text string given an image. We call this probabilistic score the Visual Generative Pre-Training Score (VisualGPTScore). While the VisualGPTScore produces near-perfect accuracy on some retrieval benchmarks, it yields poor accuracy on others. We analyze this behavior through a probabilistic lens, pointing out that some benchmarks inadvertently capture unnatural language distributions by creating adversarial but unlikely text captions. In fact, we demonstrate that even a "blind" language model that ignores any image evidence can sometimes outperform all prior art, reminiscent of similar challenges faced by the visual-question answering (VQA) community many years ago. We derive a probabilistic post-processing scheme that controls for the amount of linguistic bias in generative VLMs at test time without having to retrain or fine-tune the model. We show that the VisualGPTScore, when appropriately debiased, is a strong zero-shot baseline for vision-language understanding, oftentimes producing state-of-the-art accuracy.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Reasoning | Winoground | BLIP (VisualGPTScore, α-tuned) | Group Score | 16.8 | #46 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP (VisualGPTScore, α-tuned) | Image Score | 21.5 | #46 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP (VisualGPTScore, α-tuned) | Text Score | 36.5 | #46 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP (ITM) | Group Score | 13.3 | #50 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP (ITM) | Image Score | 15.8 | #50 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP (ITM) | Text Score | 35.8 | #50 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP (ITC) | Group Score | 6.5 | #80 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP (ITC) | Image Score | 9.0 | #80 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP (ITC) | Text Score | 28.0 | #80 of 114 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections