Papers › What You See is What You Read? Improving Text-Image Alignment Evaluation

What You See is What You Read? Improving Text-Image Alignment Evaluation

17 May 2023NeurIPS 2023 11arXiv:2305.10400archive 2025-07-28

Michal Yarom, Yonatan Bitton, Soravit Changpinyo, Roee Aharoni, Jonathan Herzig, Oran Lang, Eran Ofek, Idan Szpektor

Automatically determining whether a text and a corresponding image are semantically aligned is a significant challenge for vision-language models, with applications in generative text-to-image and image-to-text tasks. In this work, we study methods for automatic text-image alignment evaluation. We first introduce SeeTRUE: a comprehensive evaluation set, spanning multiple datasets from both text-to-image and image-to-text generation tasks, with human judgements for whether a given text-image pair is semantically aligned. We then describe two automatic methods to determine alignment: the first involving a pipeline based on question generation and visual question answering models, and the second employing an end-to-end classification approach by finetuning multimodal pretrained models. Both methods surpass prior approaches in various text-image alignment tasks, with significant improvements in challenging cases that involve complex composition or unnatural images. Finally, we demonstrate how our approaches can localize specific misalignments between an image and a given text, and how they can be used to automatically re-rank candidates in text-to-image generation.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

yonatanbitton/wysiwyr officialpytorchApache-2.0 report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Image GenerationImage to textQuestion AnsweringQuestion GenerationQuestion-GenerationText GenerationText to Image GenerationText-to-Image GenerationVisual Question AnsweringVisual Reasoning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Reasoning Winoground VQ2 Group Score 30.5 #12 of 114 Archive leaderboard report
Visual Reasoning Winoground VQ2 Image Score 42.2 #12 of 114 Archive leaderboard report
Visual Reasoning Winoground VQ2 Text Score 47 #12 of 114 Archive leaderboard report
Visual Reasoning Winoground PaLI (ft SNLI-VE + Synthetic Data) Group Score 28.75 #15 of 114 Archive leaderboard report
Visual Reasoning Winoground PaLI (ft SNLI-VE + Synthetic Data) Image Score 38 #15 of 114 Archive leaderboard report
Visual Reasoning Winoground PaLI (ft SNLI-VE + Synthetic Data) Text Score 46.5 #15 of 114 Archive leaderboard report
Visual Reasoning Winoground PaLI (ft SNLI-VE) Group Score 28.70 #18 of 114 Archive leaderboard report
Visual Reasoning Winoground PaLI (ft SNLI-VE) Image Score 41.50 #18 of 114 Archive leaderboard report
Visual Reasoning Winoground PaLI (ft SNLI-VE) Text Score 45.00 #18 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP2 (ft COCO) Group Score 23.50 #22 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP2 (ft COCO) Image Score 26.00 #22 of 114 Archive leaderboard report
Visual Reasoning Winoground BLIP2 (ft COCO) Text Score 44.00 #22 of 114 Archive leaderboard report
Visual Reasoning Winoground COCA ViT-L14 (f.t on COCO) Group Score 8.25 #78 of 114 Archive leaderboard report
Visual Reasoning Winoground COCA ViT-L14 (f.t on COCO) Image Score 11.50 #78 of 114 Archive leaderboard report
Visual Reasoning Winoground COCA ViT-L14 (f.t on COCO) Text Score 28.25 #78 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA large (ft SNLI-VE) Group Score 9.00 #81 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA large (ft SNLI-VE) Image Score 14.30 #81 of 114 Archive leaderboard report
Visual Reasoning Winoground OFA large (ft SNLI-VE) Text Score 27.70 #81 of 114 Archive leaderboard report
Visual Reasoning Winoground CLIP RN50x64 Group Score 10.25 #83 of 114 Archive leaderboard report
Visual Reasoning Winoground CLIP RN50x64 Image Score 13.75 #83 of 114 Archive leaderboard report
Visual Reasoning Winoground CLIP RN50x64 Text Score 26.50 #83 of 114 Archive leaderboard report
Visual Reasoning Winoground TIFA Group Score 11.30 #103 of 114 Archive leaderboard report
Visual Reasoning Winoground TIFA Image Score 12.50 #103 of 114 Archive leaderboard report
Visual Reasoning Winoground TIFA Text Score 19.00 #103 of 114 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections