Papers › What You See is What You Read? Improving Text-Image Alignment Evaluation
What You See is What You Read? Improving Text-Image Alignment Evaluation
Michal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni, Jonathan Herzig, Oran Lang, Eran Ofek, Idan Szpektor
Automatically determining whether a text and a corresponding image are semantically aligned is a significant challenge for vision-language models, with applications in generative text-to-image and image-to-text tasks. In this work, we study methods for automatic text-image alignment evaluation. We first introduce SeeTRUE: a comprehensive evaluation set, spanning multiple datasets from both text-to-image and image-to-text generation tasks, with human judgements for whether a given text-image pair is semantically aligned. We then describe two automatic methods to determine alignment: the first involving a pipeline based on question generation and visual question answering models, and the second employing an end-to-end classification approach by finetuning multimodal pretrained models. Both methods surpass prior approaches in various text-image alignment tasks, with significant improvements in challenging cases that involve complex composition or unnatural images. Finally, we demonstrate how our approaches can localize specific misalignments between an image and a given text, and how they can be used to automatically re-rank candidates in text-to-image generation.
In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.
Code
Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.
Code Syntology ran Syntology
Not run by Syntology. Nothing on this page verifies that the listed code works.
Tasks
Results from the paper archive 2025-07-28
| Task | Dataset | Model | Metric | Value | Rank at snapshot | Leaderboard | Report |
|---|---|---|---|---|---|---|---|
| Visual Reasoning | Winoground | VQ2 | Group Score | 30.5 | #12 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | VQ2 | Image Score | 42.2 | #12 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | VQ2 | Text Score | 47 | #12 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | PaLI (ft SNLI-VE + Synthetic Data) | Group Score | 28.75 | #15 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | PaLI (ft SNLI-VE + Synthetic Data) | Image Score | 38 | #15 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | PaLI (ft SNLI-VE + Synthetic Data) | Text Score | 46.5 | #15 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | PaLI (ft SNLI-VE) | Group Score | 28.70 | #18 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | PaLI (ft SNLI-VE) | Image Score | 41.50 | #18 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | PaLI (ft SNLI-VE) | Text Score | 45.00 | #18 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP2 (ft COCO) | Group Score | 23.50 | #22 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP2 (ft COCO) | Image Score | 26.00 | #22 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | BLIP2 (ft COCO) | Text Score | 44.00 | #22 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | COCA ViT-L14 (f.t on COCO) | Group Score | 8.25 | #78 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | COCA ViT-L14 (f.t on COCO) | Image Score | 11.50 | #78 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | COCA ViT-L14 (f.t on COCO) | Text Score | 28.25 | #78 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA large (ft SNLI-VE) | Group Score | 9.00 | #81 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA large (ft SNLI-VE) | Image Score | 14.30 | #81 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | OFA large (ft SNLI-VE) | Text Score | 27.70 | #81 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | CLIP RN50x64 | Group Score | 10.25 | #83 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | CLIP RN50x64 | Image Score | 13.75 | #83 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | CLIP RN50x64 | Text Score | 26.50 | #83 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | TIFA | Group Score | 11.30 | #103 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | TIFA | Image Score | 12.50 | #103 of 114 | Archive leaderboard | report |
| Visual Reasoning | Winoground | TIFA | Text Score | 19.00 | #103 of 114 | Archive leaderboard | report |
Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections