Datasets › GenEval

GenEval

Introduced by Dhruba Ghosh et al. in GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment17 Oct 2023 archive 2025-07-28

Recent breakthroughs in diffusion models, multimodal pretraining, and efficient finetuning have led to an explosion of text-to-image generative models. Given human evaluation is expensive and difficult to scale, automated methods are critical for evaluating the increasingly large number of new models. However, most current automated evaluation metrics like FID or CLIPScore only offer a holistic measure of image quality or image-text alignment, and are unsuited for fine-grained or instance-level analysis. In this paper, we introduce GenEval, an object-focused framework to evaluate compositional image properties such as object co-occurrence, position, count, and color. We show that current object detection models can be leveraged to evaluate text-to-image models on a variety of generation tasks with strong human agreement, and that other discriminative vision models can be linked to this pipeline to further verify properties like object color. We then evaluate several open-source text-to-image models and analyze their relative generative capabilities on our benchmark. We find that recent models demonstrate significant improvement on these tasks, though they are still lacking in complex capabilities such as spatial relations and attribute binding. Finally, we demonstrate how GenEval might be used to help discover existing failure modes, in order to inform development of the next generation of text-to-image models. Our code to run the GenEval framework is publicly available at this https URL.

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Text-to-Image Generation GenEval SD3.5-Medium+Flow-GRPO Overall 0.95 Flow-GRPO: Training Flow Matching Models via Online RL yifan123/flow_grpo 20 Compare

Papers archive 2025-07-28

16 shown of 16 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 100. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation 2 2 3 Jun 2025 not harvested
MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO 1 1 19 May 2025 not harvested
Flow-GRPO: Training Flow Matching Models via Online RL 1 1 8 May 2025 ran 1 of 9 samples (8 unverified)
Transfer between Modalities with MetaQueries 0 1 8 Apr 2025 not harvested
Lumina-Image 2.0: A Unified and Efficient Image Generative Framework 1 1 27 Mar 2025 ran 7 of 16 samples (9 unverified)
DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers 0 1 18 Mar 2025 not harvested
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer 1 2 30 Jan 2025 not harvested
Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling 1 2 29 Jan 2025 not harvested
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step 2 2 23 Jan 2025 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training 0 1 12 Dec 2024 not harvested
JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation 1 1 12 Nov 2024 not harvested
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens 1 1 17 Oct 2024 not harvested
Emu3: Next-Token Prediction is All You Need 2 1 27 Sep 2024 ran 3 of 6 samples (3 unverified)
Show-o: One Single Transformer to Unify Multimodal Understanding and Generation 1 1 22 Aug 2024 not harvested
PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation 2 1 7 Mar 2024 ran 6 of 9 samples (3 unverified; 2 pointer-only for licence)
PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models 1 1 10 Jan 2024 ran 5 of 5 samples (0 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

MIT license

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • GenEval

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections