{"url":"/dataset/geneval","name":"GenEval","full_name":null,"description_markdown":"Recent breakthroughs in diffusion models, multimodal pretraining, and efficient finetuning have led to an explosion of text-to-image generative models. Given human evaluation is expensive and difficult to scale, automated methods are critical for evaluating the increasingly large number of new models. However, most current automated evaluation metrics like FID or CLIPScore only offer a holistic measure of image quality or image-text alignment, and are unsuited for fine-grained or instance-level analysis. In this paper, we introduce GenEval, an object-focused framework to evaluate compositional image properties such as object co-occurrence, position, count, and color. We show that current object detection models can be leveraged to evaluate text-to-image models on a variety of generation tasks with strong human agreement, and that other discriminative vision models can be linked to this pipeline to further verify properties like object color. We then evaluate several open-source text-to-image models and analyze their relative generative capabilities on our benchmark. We find that recent models demonstrate significant improvement on these tasks, though they are still lacking in complex capabilities such as spatial relations and attribute binding. Finally, we demonstrate how GenEval might be used to help discover existing failure modes, in order to inform development of the next generation of text-to-image models. Our code to run the GenEval framework is publicly available at this https URL.","description_withheld":null,"homepage":"https://arxiv.org/abs/2310.11513","introduced_date":"2023-10-17","introduced_date_note":null,"introduced_by":{"paper":"/paper/geneval-an-object-focused-framework-for","title":"GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment","first_author":"Dhruba Ghosh","url":null},"license":{"name":"MIT license","url":null},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Text-to-Image Generation","url":"/task/text-to-image-generation","datasets_with_task":"/datasets/task/text-to-image-generation"}],"languages":[],"variants":["GenEval"],"data_loaders":[],"num_papers_in_archive":100,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/text-to-image-generation-on-geneval","task":"Text-to-Image Generation","dataset_variant":"GenEval","rows":20,"metrics":["Overall","Single Obj.","Two Obj.","Color Attri.","Colors","Counting","Position"],"first_row_in_archive_order":{"model":"SD3.5-Medium+Flow-GRPO","paper":"/paper/flow-grpo-training-flow-matching-models-via","metrics":{"Overall":"0.95"},"code_links":[{"title":"yifan123/flow_grpo","url":"https://github.com/yifan123/flow_grpo"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/uniworld-v1-high-resolution-semantic-encoders","title":"UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation","date":"2025-06-03","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/mindomni-unleashing-reasoning-generation-in","title":"MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO","date":"2025-05-19","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/flow-grpo-training-flow-matching-models-via","title":"Flow-GRPO: Training Flow Matching Models via Online RL","date":"2025-05-08","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":1,"samples_unverified":8,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/transfer-between-modalities-with-metaqueries","title":"Transfer between Modalities with MetaQueries","date":"2025-04-08","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/lumina-image-2-0-a-unified-and-efficient","title":"Lumina-Image 2.0: A Unified and Efficient Image Generative Framework","date":"2025-03-27","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":16,"samples_ran":7,"samples_unverified":9,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/diffmoe-dynamic-token-selection-for-scalable","title":"DiffMoE: Dynamic Token Selection for Scalable Diffusion Transformers","date":"2025-03-18","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/sana-1-5-efficient-scaling-of-training-time","title":"SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer","date":"2025-01-30","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/janus-pro-unified-multimodal-understanding","title":"Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling","date":"2025-01-29","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/can-we-generate-images-with-cot-let-s-verify","title":"Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step","date":"2025-01-23","rows_on_this_dataset":2,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":3,"samples_unverified":1,"pointer_only_for_licence":4,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/snapgen-taming-high-resolution-text-to-image","title":"SnapGen: Taming High-Resolution Text-to-Image Models for Mobile Devices with Efficient Architectures and Training","date":"2024-12-12","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/janusflow-harmonizing-autoregression-and","title":"JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation","date":"2024-11-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/fluid-scaling-autoregressive-text-to-image","title":"Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens","date":"2024-10-17","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/emu3-next-token-prediction-is-all-you-need","title":"Emu3: Next-Token Prediction is All You Need","date":"2024-09-27","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":3,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/show-o-one-single-transformer-to-unify","title":"Show-o: One Single Transformer to Unify Multimodal Understanding and Generation","date":"2024-08-22","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/pixart-s-weak-to-strong-training-of-diffusion","title":"PixArt-Σ: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation","date":"2024-03-07","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":6,"samples_unverified":3,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/pixart-d-fast-and-controllable-image","title":"PIXART-δ: Fast and Controllable Image Generation with Latent Consistency Models","date":"2024-01-10","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":5,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":6,"samples_harvested":49,"samples_ran":25,"samples_unverified":24,"pointer_only_for_licence":6,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}