Papers › SelfEval: Leveraging the discriminative nature of generative models for evaluation

SelfEval: Leveraging the discriminative nature of generative models for evaluation

17 Nov 2023arXiv:2311.10708archive 2025-07-28

Sai Saketh Rambhatla, Ishan Misra

We present an automated way to evaluate the text alignment of text-to-image generative diffusion models using standard image-text recognition datasets. Our method, called SelfEval, uses the generative model to compute the likelihood of real images given text prompts, and the likelihood can be used to perform recognition tasks with the generative model. We evaluate generative models on standard datasets created for multimodal text-image discriminative learning and assess fine-grained aspects of their performance: attribute binding, color recognition, counting, shape recognition, spatial understanding. Existing automated metrics rely on an external pretrained model like CLIP (VLMs) or LLMs, and are sensitive to the exact pretrained model and its limitations. SelfEval sidesteps these issues, and to the best of our knowledge, is the first automated metric to show a high degree of agreement for measuring text-faithfulness with the gold-standard human evaluations across multiple generative models, benchmarks and evaluation metrics. SelfEval also reveals that generative models showcase competitive recognition performance on challenging tasks such as Winoground image-score compared to discriminative models. We hope SelfEval enables easy and reliable automated evaluation for diffusion models.

PaperPDF

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

AttributeVisual Reasoning

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Visual Reasoning Winoground OCLIP (ViT-H/14) Image Score 12.75 #64 of 114 Archive leaderboard report
Visual Reasoning Winoground OCLIP (ViT-H/14) Text Score 30.75 #64 of 114 Archive leaderboard report
Visual Reasoning Winoground CLIP (ViT-L/14) Image Score 8.0 #68 of 114 Archive leaderboard report
Visual Reasoning Winoground CLIP (ViT-L/14) Text Score 30.25 #68 of 114 Archive leaderboard report
Visual Reasoning Winoground LDM-T5 (SelfEval) Image Score 13.50 #75 of 114 Archive leaderboard report
Visual Reasoning Winoground LDM-T5 (SelfEval) Text Score 29.00 #75 of 114 Archive leaderboard report
Visual Reasoning Winoground PDM-T5 (SelfEval) Image Score 12.00 #77 of 114 Archive leaderboard report
Visual Reasoning Winoground PDM-T5 (SelfEval) Text Score 28.25 #77 of 114 Archive leaderboard report
Visual Reasoning Winoground LDM-CLIP (SelfEval) Image Score 7.25 #95 of 114 Archive leaderboard report
Visual Reasoning Winoground LDM-CLIP (SelfEval) Text Score 22.75 #95 of 114 Archive leaderboard report
Visual Reasoning Winoground PDM-CLIP (SelfEval) Image Score 14.00 #107 of 114 Archive leaderboard report
Visual Reasoning Winoground PDM-CLIP (SelfEval) Text Score 17.00 #107 of 114 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Diffusion

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections